
This paper presents the design of a single chip adaptive beamformer which contains 5 million transistors and can perform 50 GigaFlops. The core processor of the adaptive beamformer is a QR-array processor implemented on a fully efficient linear systolic architecture. The paper highlights a number of rapid design techniques that have been used to realise the design. These include an architecture synthesis tool for quickly developing the circuit architecture and the utilisation of a library of parameterisable silicon intellectual property (IP) cores, to rapidly develop the circuit layouts.
A combination of linear and nonlinear methods for feature fusion is introduced and the performance of this methodology is illustrated on a real-world problem: the detection of sudden and non-anticipated lapses of attention in car drivers due to drowsiness. To achieve this, signals coming from heterogeneous sources are processed, namely the brain electric activity, variation in the pupil size, and eye and eyelid movements. For all the signals considered, the features are extracted both in the spectral domain and in state space. Linear features are obtained by the modified periodogram, whereas the nonlinear features are based on the recently introduced method of delay vector variance (DVV). The decision process based on such fused features is achieved by support vector machines (SVM) and learning vector quantization (LVQ) neural networks. For the latter also methods of metrics adaptation in the input space are applied. The parameters of all utilized algorithms are optimized empirically in order to gain maximal classification accuracy. It is also shown that metrics adaptation by weighting the input features can improve the classification accuracy, but only to a limited extent. Limited improvements are also obtained when fusing features of selected signals, but highest improvements are gained by fusion of features of all available signals. In this case test errors are reduced down to 9% in the mean, which clearly illustrates the potential of our methodology to establish a reference standard of drowsiness and microsleep detection devices for future online driver monitoring.
The process of DNA sequence matching and database search is one of the major problems of the bioinformatics community. Major scientific efforts to address this problem have provided algorithms and software tools for molecular biologists since the early 1970s. At the algorithmic and software level BLAST is by far the most popular tool. It has been developed and continues to be maintained and distributed by the NCBI organization. The BLAST algorithm and software is computationally very intensive and as a result several computer vendors use it as a benchmark. On the other hand no systematic approach for hardware speedup of BLAST and its variants for different query and database size has been reported to date. In this paper we present our architecture that implements the BLAST algorithm for all of its major versions, and for any size of database and query. The system has been fully designed and partially implemented with reconfigurable logic. It consists of software and hardware parts and achieves a speedup of several times up to thousands of times vs general purpose computers.
The energy of brain potentials evoked during processing of visual stimuli is considered as a new biometric. In particular, we propose several advances in the feature extraction and classification stages. This is achieved by performing spatial data/sensor fusion, whereby the component relevance is investigated by selecting maximum informative ( EEG) electrodes ( channels) selected by Davies-Bouldin index. For convenience and ease of cognitive processing, in the experiments, simple black and white drawings of common objects are used as visual stimuli. In the classification stage, the Elman neural network is employed to classify the generated EEG energy features. Simulations are conducted by using the hold-out classification strategy on an ensemble of 1,600 raw EEG signals, and 35 maximum informative channels achieved the maximum recognition rate of 98.56 +/- 1.87%. Overall, this study indicates the enormous potential of the EEG biometrics, especially due to its robustness against fraud.
Many evolving video services and applications for intelligent security systems require reliable transmission of high quality video to diverse clients over heterogeneous networks using available system resources. Scalable video coding (SVC) is one of the emerging video compression technologies with such potential capabilities. Advances in lifting-based motion-compensated temporal filtering (MCTF) have enabled highly efficient and flexible spatial, temporal, signal-to-noise ratio (SNR), and complexity scalability to be realized over a wide range of bit rates. In this paper, we present an algorithm to improve the update step of MCTF, which serves as an important informative step for the coding performance of SVC. A novel update-step algorithm, which takes advantage of the chrominance information of the video sequence and the correlation of the motion vectors (MVs) of the neighboring blocks as well as the correlation of the derived update MVs in the low-pass frames, is proposed to improve update step of MCTF by (1) computing correct update motion information, (2) generating correct amount of energy contained in the high-pass frames. Experimental results show that the proposed algorithm can significantly improve the quality of the reconstructed video sequence in visual quality.
Posture analysis is an active research area in computer vision for applications such as home care and security monitoring. This paper describes the design of a system for posture analysis with hardware acceleration, addressing the following four aspects: (a) a design workflow for posture analysis based on radial shape and projection histogram representations; (b) the implementation of different architectures based on a high-level hardware design approach with support for automating transformations to improve parallelism and resource optimisation; (c) accuracy evaluation of the proposed posture analysis system, and (d) performance evaluation for the derived designs. One of the designs, which targets a Xilinx XC2V6000 FPGA at 90.2 MHz, is able to perform posture analysis at a rate of 1,164 frames per second with a frame size of 320 by 240 pixels. It represents 3.5 times speedup over optimised software running on a 2.4 GHz AMD Athlon 64 3700+ computer. The frame rate is well above that of real-time video, which enables the sharing of the FPGA among multiple video sources.
This work describes the VHDL design and implementation of block-based motion estimation in order to make it feasible for real-time video applications. The design was functionally tested and simulated using ModelSim from Mentor Graphics tools, and then verified using both a VHDL testbench and the Matlab® Image processing tools. The design was tested for different image sizes at different clock frequencies with varying block sizes and search areas. With a clock frequency of 400 MHz, the estimated time for motion estimation for QCIF and CIF sequences shows the feasibility for real-time video-codec.
Tracking people across multiple cameras is a challenging research area in visual computing, especially when these cameras have non-overlapping field of views. The important task is to associate a current subject with other prior appearances of the same subject across time and space in a camera network. Several known techniques rely on Bayesian approaches to perform the matching task. However, these approaches do not scale well when the dimension of the problem increases; e.g. when the number of subject or possible path increases. The aim of this paper is to propose a unified tracking framework using particle filters to efficiently switch between visual tracking (field of view tracking) and track prediction (non-overlapping region tracking). The particle filter tracking system utilizes a map (known environment) to assist the tracking process when targets leave the field of view of any camera. We implemented and tested this tracking approach in an in-house multiple cameras system as well as using on-line data. Promising results were obtained which suggested the feasibility of such an approach.
Dynamic Voltage Scaling (DVS) is one of the techniques used to obtain energy-saving in real-time DSP systems. In many DSP systems, some tasks contain conditional instructions that have different execution times for different inputs. Due to the uncertainties in execution time of these tasks, this paper models each varied execution time as a probabilistic random variable and solves the Voltage Assignment with Probability (VAP) Problem. VAP problem involves finding a voltage level to be used for each node of an date flow graph (DFG) in uniprocessor and multiprocessor DSP systems. This paper proposes two optimal algorithms, one for uniprocessor and one for multiprocessor DSP systems, to minimize the expected total energy consumption while satisfying the timing constraint with a guaranteed confidence probability. The experimental results show that our approach achieves significant energy saving than previous work. For example, our algorithm for multiprocessor achieves an average improvement of 56.1% on total energy-saving with 0.80 probability satisfying timing constraint.
A scheme for reducing the hardware resources to implement on LUT-based FPGA devices the twiddle factors required in Fast Fourier Transform (FFT) processors is presented. The proposed scheme reduces the number of embedded block RAM for large FFTs and the number of slices for FFT lengths higher than 128 points. Results are given for Xilinx devices, but they can be generalized for other advanced LUT-based devices like ALTERA Stratix.
For hands-free communication system, this paper describes a noise reduction method using a 2-channel microphone. Recently, the Complex Spectrum Circle Centroid (CSCC) method has been proposed. This method utilizes geometric information and estimates the spectrum of the target signal. The method is advantageous in that no adjustment of the array-processing parameters to the environment is necessary before its operation and it is effective with non-stationary noise. However, the original CSCC method requires at least three microphones to estimate the spectrum of the target signal (center of circle). In this paper, we propose a method which estimates the spectrum of the target signal using only two microphones. In experimental results, the proposed method outperforms the Delay-and-Sum approach and can restore the target signal almost completely in a simulated noisy environment.
Application-specific instruction-set processors (ASIPs) provide a good alternative for video processing acceleration, but the productivity gap implied by such a new technology may prevent leveraging it fully. Video processing SoCs need flexibility that is not available in pure hardware architectures, while pure software solutions do not meet video processing performance constraints. Thus, ASIP design could offer a good tradeoff between performance and flexibility. Video processing algorithms are often characterized by intrinsic parallelism that can be accelerated by ASIP specialized instructions. In this paper, we propose a new approach for exploiting sequences of tightly coupled specialized instructions in ASIP design applicable to video processing. Our approach, which avoids costly data communications by applying data grouping and data reuse, consists of accelerating an algorithm’s critical loops by transforming them according to a new intermediate representation. This representation is optimized and loop parallelism possibilities are also explored. This approach has been applied to video processing algorithms such as the ELA deinterlacer and the 2D-DCT. Experimental results show speedups up to 18 (on the considered applications, while the hardware overhead in terms of additional logic gates was found to be between 18 and 59%.
Occlusion is a difficult problem for visual tracking and we use multiple wide baseline cameras to deal with occlusion. We propose a data fusion approach for visual tracking using multiple cameras with overlapping fields of view. First, we present a spatial and temporal recursive Bayesian filter to fuse information from multiple cameras. An adaptive particle filter is formulated to realize the spatial and temporal recursive Bayesian filter. Our algorithm is able to recover the target's position even under complete occlusion in a camera.
We present a hierarchical architecture and learning algorithm for visual recognition and inference tasks such as imagination, reconstruction of occluded images, and expectation-driven segmentation. Certain characteristics of biological vision are used for guidance, such as extensive feedback and lateral recurrence, a highly overcomplete early stage (VI) and sparse distributed activity. Recent advances in computational methods for learning overcomplete dictionaries are used to explore how overcompleteness can be useful for visual tasks. We posit a stochastic, hierarchical generative-world-model (GWM) and develop a simplified-world-model (SWM) based on a variational approximation to the Boltzmann-like distribution. The SWM is designed to enforce sparsity and leads to a tractable dynamic network. Experimentally, we show that increasing the degree of overcompleteness results in improved recognition and segmentation. Critical to the success of this vision system is the sparse coding of images using a learned overcomplete dictionary. An algorithm for performing dictionary learning termed FOCUSS-CNDL is developed in Chapter 2. In tests with natural images, learned overcomplete dictionaries are shown to have higher coding efficiency than complete dictionaries: images encoded with an overcomplete dictionary have both higher compression (fewer bits/pixel) and higher accuracy (lower mean-square error). The vision algorithm of Chapter 1 requires non-negative sparse codes, which is discussed in Chapter 3. A non-negative version of the FOCUSS algorithm is shown to be superior to a matching-pursuit variant. Also, the FOCUSS-CNDL algorithm is found to have better image coding performance than another overcomplete independent analysis (ICA) algorithm. The final chapter presents methods for detecting rare events in a time series of noisy and nonparametrically-distributed data. These algorithms are tested on a difficult real-world problem: predicting failures in hard-drives. An algorithm is developed based on the multiple-instance learning framework and the naive Bayesian classifier (mi-NB) which is specifically designed for the low false-alarm case. Other methods compared are support vector machines (SVMs), unsupervised clustering, and non-parametric statistical tests. While not specific to vision tasks, the mi-NB algorithm may find uses in semi-supervised image categorization tasks.
H.264/AVC also known as MPEG 4 part 10 or JVT, is a recently established video coding standard by the Joint Video Team (JVT) of the ISO/IEC MPEG and ITU-T VCEG. The main goal of the paper is to give a broader understanding of the design considerations for the transform and quantization blocks from H.264/AVC, by presenting area and speed optimized implementations of these blocks. The area optimized design can be used in low performance applications like mobile devices, while the speed optimized designs can be used in high definition encoders. Various designs with these blocks were synthesized with 0.18 μm TSCM technology and were also implemented on a Xilinx FPGA. The resulting gate counts were anywhere from 294 to 47,762 gates and the throughput was anywhere from 6 to 2,552 M pixels/s depending on block and optimization. In addition, a system on a programmable chip implementation of the DCT and quantization blocks is presented, which uses the Xilinx Virtex II-Pro's FPGA and its Power PC. Using this system it is possible to process 0.8 M pixels/s.
With the prevalence of video-on-demand (VOD) services as well as the diffusion of various multimedia devices, caching in a multimedia streaming server is becoming increasingly important. However, due to some peculiar characteristics of multimedia objects and user activities in streaming services, design of an efficient caching system becomes a more challenging problem compared to the traditional caching systems. This paper discusses some important issues that are of interest in the domain of multimedia streaming caching and presents a new cache management scheme for multimedia streaming servers. Our new scheme considers different streaming rates of multimedia objects as well as the inter-arrival time between two consecutive requests on an identical object. It also considers user activities in requesting and playing multimedia contents. Trace-driven simulations with real world VOD traces show that the proposed scheme improves the performance of multimedia streaming systems significantly.
Many discriminative classification algorithms are designed for tasks where samples can be represented by fixed-length vectors. However, many examples in the fields of text processing, computational biology and speech recognition are best represented as variable-length sequences of vectors. Although several dynamic kernels have been proposed for mapping sequences of discrete observations into fixed-dimensional feature-spaces, few kernels exist for sequences of continuous observations. This paper introduces continuous rational kernels, an extension of standard rational kernels, as a general framework for classifying sequences of continuous observations. In addition to allowing new task-dependent kernels to be defined, continuous rational kernels allow existing continuous dynamic kernels, such as Fisher and generative kernels, to be calculated using standard weighted finite-state transducer algorithms. Preliminary results on both a large vocabulary continuous speech recognition (LVCSR) task and the TIMIT database are presented.
This paper presents a novel dynamic and scalable caching algorithm of proxy server with a finite storage size for multimedia objects. Among the multimedia such as text, image, audio and video, video is a dominant component in terms of the performance of proxy server due to its traffic characteristics. For the fast caching process, caching sequences for videos are obtained to decrease both the buffer size and the required bandwidth and saved into metafiles in advance. Then, we present a novel caching and replacing algorithms for multimedia objects based on the metafiles. Finally, experimental results are provided to show the superior performance of the proposed algorithm.