
The interest in this paper is in efficient configuration of automatic speech recognition (ASR) systems for use by under-served speaker populations. A task domain involving Indian farmers accessing information on agricultural commodities through a spoken dialog system in multiple languages is presented. To facilitate the development of ASR system for this domain, a speech corpus was collected in rural areas from speakers of four languages over wireless cellular channels. This paper investigates the problem of ASR acoustic modelling for this task domain. Continuous density hidden Markov model (CDHMM) and subspace Gaussian mixture model (SGMM) [1] based techniques are used to train acoustic models in four languages: Assamese, Bengali, Hindi and Marathi. Issues relating to limited linguistic resources with their impact on ASR word accuracy for these languages are addressed.
In Cognitive Radio Network, secondary users opportunistically use the licensed channels to transmit data. The secondary users first conduct the local sensing to determine the unoccupied channels and then choose the satisfied channel in the available channel set to transmit. However, the sensing results might not be accurate due to many problems: channel fading or shadowing, noise, hidden primary users or the aggregated interference caused by other secondary user transmissions. In this paper, we are interested in improving the spectrum sensing performance by reducing the missed-detection sensing caused by the hidden primary user problem. Particularly, we first proposed the channel state transition diagram capturing the hidden terminal possibility and then developed a collaborative scheme to exchange the sensing result between secondary users. Based on the summarized sensing results achieved from secondary user neighbors and local sensing, each secondary will build its own channel table listing all available channels. Numerical results show that the proposed spectrum sensing scheme enhances the sensing results and the network performance.
Speech enhancement problem using hidden Markov model (HMM) and minimum mean square error (MMSE) in cepstral domain is studied. This noise reduction approach can be considered as weighted-sum filtering of the noisy speech signal in which the filters weights are estimated using the HMM of noisy speech. To have an accurate estimation of the noisy speech HMM, vector Taylor series (VTS) is proposed and compared with the parallel model combination (PMC) technique. Furthermore, proposed cepstral-domain HMM-based speech enhancement systems are compared with the renowned autoregressive HMM (AR-HMM) approach. The evaluation results confirm the superiority of the cepstral domain approach in comparison with AR-HMM and indicate that VTS-based method provides higher noise reduction rate than the PMC-based method.
In recent years there has been an increasing interest in speech and language processing systems dedicated to Arabic language. In order to perform adequate design and evaluation of those systems, speech databases are needed. The aim of this paper is to evaluate the design of Arabic and English speech recognition systems by using common acoustic models. Cross-language experiments between Arabic and English are conducted and discussed with respect to the main class of phonemes in each language. The LDC WestPoint Arabic database and TIMIT are used in these experiments. The results show that lack of enough speech resource that faces Arabic language can be solved by considering models' features of common phonemes given by English.
The detection and characterization of burst signals are challenging tasks for time-frequency analysis, due to their very short duration. This paper investigates in this context the recurrence plot analysis (RPA) method, from which it derives the vector samples processing (VeSP) concept. The paper shows that VeSP is a generic framework that unifies signal processing concepts like histogram and autocorrelation, which it also generalizes and extends. Results of VeSP based tools are provided, concerning detection of transient signals, noise reduction, and frequency estimation.
In Digital Beamforming (DBF) radars, multiple beams are formed simultaneously at different azimuth and/or elevation angles to monitor large areas of interest. In order to estimate the location of a target in azimuth without specifically steering the beams at the target or using special techniques such as monopulse, simple centroid schemes can be used. This assumes that a target is seen by more than one beam, which requires the beam pattern to be somewhat overlapping and that the Signal-to-Noise Ratio (SNR) of the target is sufficient. In this paper, a new azimuth centroid computation scheme using amplitude comparison is presented. Since the digital phased array beam shape is different at each azimuth, each beam is modeled by simple polynomial curve fitting. Simulation and preliminary experimental results show that the proposed method gives better azimuth accuracies compared to other centroid techniques.
This communication summarizes the outcome of our research program on the design of a diagnostic system for neuromuscular disorders based on the analysis of human movement using the Kinematic Theory of Rapid Human Movements. Herein, this design problem is split in sub-problems which are then described. The solutions adopted at each design step are explained. As an example of application, typical results obtained so far for the assessment of the most important modifiable risk factors of brain stroke (diabetes, hypertension, hypercholesterolemia, obesity, cardiac problems, and cigarette smoking) are reported by the means of the area under the receiver operating characteristic curve (AUC).
Multi-target filtering aims at tracking an unknown number of targets from a set of observations. The Probability Hypothesis Density (PHD) Filter is a promising solution but cannot be implemented exactly. Suboptimal implementation techniques include Gaussian Mixture (GM) solutions, which hold only in linear and Gaussian models, and Sequential Monte Carlo (SMC) algorithms, which estimate the number of targets and their state parameters for a more general class of models. In this paper, we address the case of Gaussian models where the state can be decomposed into a linear component and a non-linear one, and we show that the use of SMC methods in such models can indeed be reduced. Our technique not only improves the estimate of the number of targets but also that of their state. We finally adapt the technique to linear and Gaussian jump Markov state space systems (JMSS) in order to reduce the intractability of existing solutions, and to JMSS with partially linear and partially non-linear state vector.
The main challenge of ground penetrating radar (GPR) based land mine detection is to have an accurate image analysis method that is capable of reducing false alarms. However an accurate image relies on having sufficient spatial resolution in the received signal. But because the diameter of an AP mine can be as low as 2cm and many soils have very high attenuations at frequencies above 3GHz, the accurate detection of landmines is accomplished using advanced algorithms. Using image reconstruction and by carrying out the system level analysis of the issues involved with recognition of landmines allows the landmine detection problem to be solved. The SIMCA ('SIMulated Correlation Algorithm') is a novel and accurate landmine detection tool that carries out correlation between a simulated GPR trace and a clutter1 removed original GPR trace. This correlation is performed using the MATLAB® processing environment. The authors tried using convolution and correlation. But in this paper the correlated results are presented because they produced better results. Validation of the results from the algorithm was done by an expert GPR user and 4 other general users who predict the location of landmines. These predicted results are compared with the ground truth data.
The aim of this paper is to investigate the influence of Adaptative Multi-Rate Wideband (AMR-WB) speech coding on Distributed Speaker Recognition (DSR). The main goal is to improve speaker recognition performance without resynthesizing the speech waveform in mobile communications. For this purpose, we have implemented a method in order to extract the acoustic features in the compressed domain, directly from encoded bitstream. The obtained results, using ARADIGIT database, show that the proposed approach, with GMM and SVM models, using ISF (Immittance Spectral Frequency) parameters through a noised channel (AWGN and Rayleigh) are promising.
We present a summary of our recent work on using lattice-form Mach-Zehnder interferometers implemented using silica-based planar lightwave circuit technology for optical signal processing. In particular, we demonstrate that the same device structure (either based on 6 taps or 12 taps) can be used to perform various signal processing functions ranging from pulse repetition rate multiplication to arbitrary waveform generation.
We describe a learning-based method for recovering 3D human body pose from single images and monocular image sequences. Our approach requires neither an explicit body model nor prior labeling of body parts in the image. Instead, it recovers pose by direct nonlinear regression against shape descriptor vectors extracted automatically from image silhouettes. For robustness against local silhouette segmentation errors, silhouette shape is encoded by histogram-of-shape-contexts descriptors. We evaluate several different regression methods: ridge regression, Relevance Vector Machine (RVM) regression, and Support Vector Machine (SVM) regression over both linear and kernel bases. The RVMs provide much sparser regressors without compromising performance, and kernel bases give a small but worthwhile improvement in performance. The loss of depth and limb labeling information often makes the recovery of 3D pose from single silhouettes ambiguous. To handle this, the method is embedded in a novel regressive tracking framework, using dynamics from the previous state estimate together with a learned regression value to disambiguate the pose. We show that the resulting system tracks long sequences stably. For realism and good generalization over a wide range of viewpoints, we train the regressors on images resynthesized from real human motion capture data. The method is demonstrated for several representations of full body pose, both quantitatively on independent but similar test data and qualitatively on real image sequences. Mean angular errors of 4-6 degrees are obtained for a variety of walking motions.
Image abstraction has wide spectrum of application, especially for computationally intensive application such as content-based image retrieval (CBIR). A recently developed abstraction system was demonstrated to enhance the performance of CBIR systems while reducing the space and time complexities. The system is based on Singular Value Decomposition (SVD), a tool borrowed from linear algebra. However, pre-abstracting images also requires high computational power, and has to be optimized for parallel environments in order to be practical. In this work we aim at optimizing the abstraction and creating a parallel engine for it. The new parallel system is demonstrated to run in less than one tenth the time required for the original system.
In the last decade, gait analysis has become one of the most active research topics in biomedical research engineering partly due to recent development of sensors and signal processing devices and more recently depth cameras. The latters can provide real-time distance measurements of moving objects. In this context, we present a new way to reconstruct body volume in motion using multiple active cameras from the depth maps they provide. A first contribution of this paper is a new and simple external camera calibration method based on several plane intersections observed with a low-cost depth camera which is experimentally validated. A second contribution consists in a body volume reconstruction method based on visual hull that is adapted and enhanced with the use of depth information. Preliminary results based on simulations are presented and compared with classical visual hull reconstruction. These results show that as little as three low-cost depth cameras can recover a more accurate 3D body shape than twenty regular cameras.
Accurate instantaneous frequency (IF) estimation of the non-stationary heart rate signal is important in quantifying the heart rate variability (HRV) measures. This study compares the effectiveness of four IF estimation methods in analyzing HRV signals. Specifically, they are the direct localization of the maximal peaks in the signal time-frequency distribution (TFD), IF estimation based on component linking technique in the TFD, IF estimation using the TFD with optimal windows based on intersection of confidence intervals rule and complex demodulation. Results of applying the IF estimation methods to synthesized and real piglet HRV signals reveal that, the approach using component linking technique outperform the other techniques with respect to the accuracy and implementation. It provides new insights in studying the evolution of the autonomic nervous regulation of the cardiovascular function over time.
In recent years, several metrics have been developed for measuring image visual quality, including the MSSIM and the visual information fidelity (VIF). However, these metrics are not robust to spatial shifts, meaning that when the reference and distorted images are misaligned by a few pixels, these metrics will produce very low scores, which is undesirable. In this paper, we extend the SSIM metric to make it robust to spatial shifts by first pre-processing the input images with the Fast Fourier transform (FFT). We then apply the magnitude of the transformed Fourier coefficients to the existing metrics because these coefficients are shift-invariant. Our assumption is that if we shift the image by a small amount of pixels, then it will not affect the perceived quality. Experimental results show that the proposed method is attractive for measuring the visual quality of 2D images as it is far less complex than the current approach, which consists in performing global motion estimation to align the input images prior to applying the metrics, and offers better accuracy.
Magnetic resonance imaging (MRI) pelvimetry measurements are useful in the diagnosis of pelvic organ prolapse given the inaccuracy of clinical examination. However, MRI measurements are currently performed manually and can be inconsistent, time-consuming and inaccurate. In this paper, we present a scheme for semi-automatic measurements on MR images based on multi scale wavelet analysis. The experiments on the MR images show that the presented scheme can detect the points of reference on the pelvic bone structure to determine the lines needed for the assessment of pelvic organ prolapse. This may lead towards more accurate and faster pelvic organ prolapse diagnosis on dynamic MR studies, and possible screening procedures for predicting predisposition to pelvic organ prolapse by radiologic evaluation of pelvimetry measurements.
The lp-norm regularized least square technique has been effectively exploited for sparse reconstruction problems. However, the choice of an optimum regularization parameter in the optimization routine still remains a challenge. In this paper we propose a new criterion which is based on MNDL, a new method for optimum subspace selection in data representation, to select the optimum regularization parameter utilizing lp-regularized least-squares. Simulations are done for combined model order selection and parameter estimation for the ubiquitous sinusoids-in-noise model. The results show that the MNDL based regularization parameter selection outperforms the state of the art methods that use MDL for the correct estimation of number of components in the signal.
A framework for development of segmentation-free optical recognizers of ancient manuscripts, which work free from line, word, and character segmentation, is proposed. The framework introduces a new representation of visual text using the concept of signature patches. These patches which are free from traditional guidelines of text, such as the baseline, are registered to each other using a microscale registration method based on the estimation of the active regions using a multilevel classifier, the directional map. Then, an one-dimensional feature vector is extracted from the registered signature patches, named spiral features. The incremental learning process is performed using a sparse representation using a dictionary of spiral feature atoms. The framework is applied to the George Washington database with promising results.