
The primary objective of the brain is to collect information from the body via nervous system and then either process those information or store them. These stored information are basically known as memory. Over the years Electroencephalography (EEG) signals have been used quite efficiently to evaluate the cognitive state of brain. In this paper, we aim to study the cognitive functioning of the brain when a user memorizes and then recognizes an object from a list of similar ones. For this purpose, we have employed wavelet transform to extract the relevant features pertaining to the two mental activities and support vector machine to distinguish between the two states. The classification accuracy, thus obtained, is above 79% for all subjects. It is also inferred from this study that during memorization the signals from the frontal and temporal lobes are dominant and during recognition, signals from the frontal and parietal lobes are dominant.
This paper proposes a mechanism that represents behavioural modelling of Indian Classical Music based on Unified Modelling Language. The raga forms the backbone of Indian music and it is the combination of several notes sequences into a composition in a way, which is pleasing to the ear. The main objective behind the work is that it can be used as a good basis for retrieval of music information of Indian Classical Music songs. Unified Modelling Language is a visual language for producing and displaying the software designs. We apply this UML representation technique, mainly Sequence Diagram, Object Diagram, and Class Diagram to represent a song.
Multiple occurrences of single object in same page or different pages of same document can generally be handled by a repository of reusable document component (RDCR). The same objects of different size and orientation could also be tagged in the same RDCR when orientation and scale maps are used as pointers. Hence after object collection and flattening in document image processing system, the compression ratio would be improved by employing OS-RDC (Orientation and Scale Mapped Reusable Document Component) which in turn would improve the performance of the system (as printing system). But especially for the scaled down and non-orthogonally rotated objects, jagging effect creates significant imaging artifacts in terms of image quality degradation. The current paper addresses to improve the perceptual image quality through OS-RDC and perception based anti-aliasing (PAA) filter. The combination of OS-RDC (used to achieve better speed) and PAA (used to improve the quality of the geometric transformed images) ensures both compression ratio and IQ (Image Quality) which are generally antagonistic in nature.
Automated control of mobile robot navigation is a challenging area in the field of robotics research. In this work, an attempt is made to use a new neural network training algorithm based on gravitational search (GS) and feed forward neural network (FFNN) for automatic robot navigation of wall following mobile robots. The GS strategy is used for setting the optimal weight set of the FFNN so as to increase the performance of the neural network. The algorithm is tested with three large datasets obtained from UCI machine learning repository, containing a sequence of sensor readings where sensors are arranged around the waist of the SCITOS G5 robot. The proposed method shows promising results for all the datasets.
In this paper, we propose a combined classifier approach based on Inner Distance Shape Context (IDSC) and Local Binary Pattern (LBP) to classify shapes accurately. The inner-distance is insensitive to shape articulations and the LBP is invariant to rotation and shift of the shape. The Dynamic Programming (DP) in case of IDSC and Earth Movers Distance (EMD) metric in case of LBP were respectively employed to obtain similarity and hence used to classify given query shape based on maximum similarity value. The experiments are conducted on publicly available shape datasets namely MPEG-7, Kimia-99, Kimia-216, Myth and Tools-2D and the results are presented by means of Bulls eye score and precision-recall metric. The comparative study is also provided with the well known approaches to determine the retrieval accuracy of the proposed approach. The experimental results demonstrate that the proposed approach yield significant improvements over baseline shape matching algorithms.
Biological vision system extracts depth from the difference in the left and right eye images. Numerous algorithms and their hardware implementations that compute disparity in real time have been proposed. However, most of them compute disparity through complicated functions that are difficult to realize in hardware and are biologically unrealistic. The brain most likely uses simpler methods to extract depth information and hence newer methodologies that could perform stereopsis with brain like elegance need to be explored. Physiological findings support the presence of disparity tuned cells in the visual cortex and show that the perception of depth evolves with experience and is not present at the time of birth. Therefore adaptively learning disparities may indeed be the algorithm underlying depth computations in the developing brain. This paper proposes a novel VLSI design using time-staggered Winner Take All to adaptively create disparity tuned cells.
A voltammetric electronic tongue has been developed for black tea assessment. While attempting to standardize the voltammetric electronic tongue a basic taste recognition test has been conducted. This electronic tongue works on cyclic voltammetric principle and has a three electrode configuration (working, reference, and counter electrode). The voltage is applied across the working and reference electrode and the response current is obtained from the counter and reference electrode. Now the huge amount of data, obtained from these response current pulses, was compressed by discrete wavelet transform (DWT). In this paper, in order to determine the optimum level of compression mean square error, separation index and principal component analysis have been employed.
Accurate predictions of traffic plays an important role in the development of an Intelligent Transport System (ITS). This paper focuses on the combination of several neural experts forming a Committee of Experts that can be used to predict traffic. Since different mechanisms have unique properties that depend on the formulae used for construction and the training data, combining the predictions of multiple experts would inevitably result in a more accurate prediction than using a single expert. The input for these experts is historic speed data, derived from vehicle-specific sensors like GPS. Furthermore, each expert within the committee is assigned a different granularity of data as its training set based on its fundamental properties, such that the expert would be most suited to making accurate predictions using that granularity of data. The different granularities of data used by experts to create models include weekwise, daywise and timewise. Each expert further recognizes the degree of influence of certain factors such as weather, holidays etc. on the flow of traffic. Eventually each expert in the committee machine makes an independent prediction of traffic and these predictions are combined using a weighted average in order to obtain the final prediction.
The paper tackles the task of preserving depth cue in single images for human perception of depth despite compression artifact. The process is mainly divided into 2 steps; saliency map computation and image compression on the basis of salient regions. The idea is that the regions which are less blurred should be more salient rather than blurred regions. For saliency computation the proposed method first obtains the blur information for the image using frequency information. Then we generate saliency map by region combining on the basis of this blur index. These saliency values are used to vary the quality parameter over the blocks in the image for JPEG compression. Experimental results show that the proposed scheme yields good results over normal JPEG compresssion in terms of PSNR values.
Over the last few years, a few researchers have made attempts to bridge the gap between Training (high-resolution gallery) and Testing (degraded, low quality probes) sets for Face Recognition under the surveillance conditions, using efficient low-level processing and statistical learning methods. In this paper, this challenging task of FR in degraded conditions have been handled using a Bag-of-Words (BOW) based approach for FR, combined with Domain Adaptation (DA). An adaptive-SIFT feature, extracted with spatially varying density at the fiducial points on the face. The SIFT features are used to form the dictionary of BOW and then combined with Local Binary Patterns (LBP). The sampling of the keypoints is denser in the discriminative parts of the face, while it is sparsely sampled at some less-interesting (pre-decided) zones of the face. An unsupervised method of DA has been used to boost the performance of FR using a BOW-based face representation. The unsupervised method of DA considers the source domain as the Training set and the target domain as the Test set. The novelty and contribution of this work on FR for surveillance applications, is the estimation of the transformation from source to target, based on the eigen-vectors of the source and the target domains using the discriminative BOW-based representation which combines two discriminative features extracted from the face. Performance analysis of the proposed method using ROC and CMC measures along with retrieval results on two real-world surveillance face datasets, shows the superiority over recent state-of-the-art techniques.
Content Based Image Retrieval (CBIR) involves the process of searching and retrieving relevant images from a database. Most CBIR systems rely on the global content of the image but the desired content in an image is often localized, demanding the need for an object-centric CBIR. We propose a cognitive inspired framework WOR (What Object to Retrieve), to solve the problem of an object-centric CBIR. The key contributions in the proposed approach are: (i) Automatic creation of class dendogram in kernel feature space, for an efficient object recognition task; (ii) Integration of information from "What" and "Where" channels in an iterative feedback mechanism, to filter erroneous contents in the outputs of individual modules. Finally, our system extracts HOG feature descriptor from the output of Where channel, for detecting similarity and rank-order the retrieved samples. Experimentations are done with real-life datasets (including PASCAL VOC) exhibits the superior performance.
In a constant endeavour to improve the quality evaluation methodology of black tea in a rapid and objective manner, electronic tongue has been a promising technology. In this direction, a multi-frequency pulse voltammetric electronic tongue is proposed towards the quantification of specific astringent compounds in tea having direct bearing on tea taste. Twenty tea samples were analysed with a voltammetric electrode array of noble metals, i.e. gold, iridium, palladium, platinum and rhodium. The electronic tongue response has been preprocessed and prediction models have been developed with partial least squares regression (PLSR) technique. The concentrations of simple theaflavin (TF1) and Theaflavin-3'-gallate (TF3) have been estimated from the electronic tongue response. The predictions of electronic tongue are in good agreement with the high performance liquid chromatography (HPLC) based estimations of TF1 and TF3. The electronic tongue predictions have been used to obtain the quality perception of tea in terms of briskness values.
The electronic tongue (ET) system is a multi-electrode system where each electrode generates a specific electronic response in presence of different chemical substances in the sample. The efficiency of an ET system mostly depends on the discriminating power of the electronic signature generated by the electrode array. In this work, a sliding window approach is used to extract discrete cosine transform (DCT) coefficients from the ET response and the corresponding energy for different position of window are used as the features of the ET response. The efficacy of the proposed method is verified on three types of ET data sets to predict four different quality of black tea using two kernel classifiers, namely support vector machine (SVM) and vector valued regularized kernel function approximation (VVRKFA).
The paper proposes an efficient human expression recognition method in transformed domain using discrete Contourlet transform (DCT). DCT represents smooth contour information in different directions reflecting human perception and so relevant to recognize facial expressions more accurately. Each face is decomposed using discrete Contourlet transform up to fourth level and coefficients of high frequency and low frequency components with varied scales and angles are obtained. Logarithmic invariant moments of directional coefficients at different levels and histogram analysis are used to build the feature vectors for classification. To reduce dimension of the feature vectors, directional subbands are selected by analyzing the entropy of the feature vectors. Support Vector Machine (SVM) is applied to classify different expressions using the proposed method. Experimental results show promising performance applied on JAFFE and Cohn-Kanade database compare to other transformed domain methods.
We present a novel Content Based Medical Image Retrieval (CBMIR) scheme for color endoscopic images using Multi-scale Geometric Analysis (MGA) of Nonsubsampled Contourlet Transform (NSCT) and the statistical framework based on Generalized Gaussian Density (GGD) model and Kullback-Leibler Distance (KLD). The subband images obtained from the NSCT decomposition are divided into number of blocks and then the coefficients of each block of each subband is modeled with GGD parameters and computing the similarity using the KLD among the model parameters. The retrieval performance of the proposed system is further improved using Least Square-Support Vector Machine (LSSVM) classifier. Extensive experiments were carried out to evaluate the effectiveness of the proposed system on endoscopic image databases consisting of 276 images. Experimental results show that the proposed CBMIR system performs efficiently in image retrieval paradigm.
This paper presents a novel approach to detect salient keypoints from the textured region in iris of an individual for efficient recognition. Salient regions are visually pre-attentive distinct portions in an image. Entropy from local segments within an image is used as the significant measure of saliency. To know the entropy value of such portions, an entropy map is generated. Local feature extraction is achieved from the segmented annular iris. Salient keypoint detection is performed using the proposed method, and subsequently each point is represented using Speeded-Up Robust Feature (SURF) descriptor. Finally, matching is performed and the recognition accuracy is calculated from the Receiver Operating Characteristic (ROC) curve. All experiments are done on publicly available BATH and CASIAv3 databases.
A novel improved Non-local Means (NLM) filtering scheme for MR images is proposed based on a simplified pulse-coupled neural network (PCNN) which follows mammal's visual/perceptual response characteristics. The proposed scheme is a two-stage filtering technique. In the first stage, the noisy MR image is passed through a PCNN and the firing time output (time signature) of the PCNN is recorded. The time signature is used to pre-select a subset of pixels in non-local vicinities, depending on their similarities with the pixel of interest. In the second stage, improved unbiased NLM (IUNLM) is applied to denoise the pixel of interest only considering the contributions of the pre-selected subset of pixels in the non-local vicinities. Extensive experiments and comparisons with state-of-the-art MRI denoising techniques were carried out to evaluate the performance of the proposed scheme. The results suggest supporting evidence of the effectiveness of the proposed scheme in MR image denoising.
We propose an approach for sparse representation of dense features for action classification. Sparse representation has already been shown in literature as a good approximation for signals for various computer vision applications. This property is leveraged to represent a dense feature like action bank in the form of sparse dictionaries. These dictionaries are learnt using on-line dictionary learning (ODL) which further facilitates incorporating new training examples into existing dictionaries for more robust representation of various categories of action as and when required. Evaluation of the proposed method on realistic action datasets like UCF50 and HMDB51 shows that considering sparse representation of a dense feature is more suitable for classification than the feature itself.
This paper presents an effective scheme to identify the abnormal mammograms in order to detect the breast cancer. The scheme utilizes the segmentation-based fractal texture analysis (SFTA) method to extract the textural features from the mammograms for the classification of normal and abnormal mammograms. A fractal analysis has been applied to collect the qualitative information of textural features. A fast correlation-based filter (FCBF) method has been used to select feature subsets containing significant features, which are used for classification purpose. The scheme was tested on the mammogram images of MIAS database. In this paper, support vector machine (SVM) has been utilized for classification of mammograms. Simulation results show an optimal classification performance index as the area under the curve (AUC) of 0.9831 in the ROC analysis.
This paper presents use of hum of a person as a biometric cue for person recognition task. Mel Frequency Cepstral Coefficients (MFCC) is found to be state-of-the-art in the voice biometrics. However, it is magnitude-based features and ignores the phase information. This paper shows the effectiveness of phase-based information extracted via Modified Group Delay Function (MODGDF). The features developed by Mel filtering of MODGDF spectrum are called Modified Group Delay Cepstral Coefficients (MGDCC). The paper demonstrates two types of fusion strategies, viz., score-level and feature-level. The experimental results show that overall performance is improved by 3 % if a score-level fusion is employed between MFCC and MGDCC and 19.78 % by feature-level fusion in terms of % Equal Error Rate (EER). These experimental results clearly indicate that incorporating phase information along with magnitude-based features can effectively captures person-specific characteristics in humming.