
Recent face recognition techniques have achieved remarkable successes in fast face retrieval on huge image datasets. But the performance is still limited when large illumination, pose, and facial expression variations are presented. In contrast, the human brain has powerful cognitive capability to recognize faces and demonstrates robustness across viewpoints, lighting conditions, even in the presence of partial occlusion. This paper proposes a closed-loop face retrieval system that combines the state-of-the-art face recognition method with the powerful cognitive function of the human brain illustrated in electroencephalography signals. The system starts with a random face image and outputs the ranking of all of the images in the database according to their similarity to the target individual. At each iteration, the single trial event related potentials (ERP) detector scores the user's interest in rapid serial visual presentation paradigm, where the presented images are selected from the computer face recognition module. When the system converges, the ERP detector further refines the lower ranking to achieve better performance. In total, 10 subjects participated in the experiment, exploring a database containing 1,854 images of 46 celebrities. Our approach outperforms existing methods with better average precision, indicating human cognitive ability complements computer face recognition and contributes to better face retrieval.
After 20 years of extensive study in psychology, some musical factors have been identified that can evoke certain kinds of emotions. However, the underlying mechanism of the relationship between music and emotion remains unanswered. This paper intends to find the genuine correlates of music emotion by exploring a systematic and quantitative framework. The task is formulated as a dimensionality reduction problem, which seeks the complete and compact feature set with intrinsic correlates for the given objectives. Since a song generally elicits more than one emotion, we explore dimensionality reduction techniques for multi-label classification. One challenging problem is that the hard label cannot represent the extent of the emotion and it is also difficult to ask the subjects to quantize their feelings. This work tries utilizing the electroencephalography (EEG) signal to solve this challenge. A learning scheme called EEG-based emotion smoothing ( E 2 S) and a bilinear multi-emotion similarity preserving embedding (BME-SPE) algorithm are proposed. We validate the effectiveness of the proposed framework on standard dataset CAL-500. Several influential correlates have been identified and the classification via those correlates has achieved good performance. We build a Chinese music dataset according to the identified correlates and find that the music from different cultures may share similar emotions.
Previous research on social interaction among humans suggested that interpersonal motor coordination can help to establish social rapport. Our research addresses the question of whether, in a human-humanoid interaction experiment, the human's overall perception of a robot can be improved by realizing motor coordination behavior that allows the robot to adapt in real-time to a person's behavior. A synchrony detection method using information distance was adopted to realize the real-time human-robot motor coordination behavior, which guided the humanoid robot to coordinate its movements to a human by measuring the behavior synchrony between the robot and the human. The feedback of the participants indicated that most of the participants preferred to interact with the humanoid robot with the adaptive motor coordination capability. The results of this proof-of-concept study suggest that the motor coordination mechanism improved humans' overall perception of the humanoid robot. Together with our previous findings, namely that humans actively coordinate their behaviors to a humanoid robot's behaviors, this study further supports the hypothesis that bidirectional motor coordination could be a valid approach to facilitate adaptive human-humanoid interaction.
In this paper, we focus on how to locate the relevant or discriminative brain regions related with external stimulus or certain mental decease, which is also called support identification, based on the neuroimaging data. The main difficulty lies in the extremely high dimensional voxel space and relatively few training samples, easily resulting in an unstable brain region discovery (or called feature selection in context of pattern recognition). When the training samples are from different centers and have between-center variations, it will be even harder to obtain a reliable and consistent result. Corresponding, we revisit our recently proposed algorithm based on stability selection and structural sparsity. It is applied to the multicenter MRI data analysis for the first time. A consistent and stable result is achieved across different centers despite the between-center data variation while many other state-of-the-art methods such as two sample t-test fail. Moreover, we have empirically showed that the performance of this algorithm is robust and insensitive to several of its key parameters. In addition, the support identification results on both functional MRI and structural MRI are interpretable and can be the potential biomarkers.
The study of cerebellum has resulted in a common agreement that it is implicated in motor learning for movement coordination. Learning governed by error signal through synaptic eligibility traces has been proposed to be a learning mechanism in cerebellum. In this paper, we extend this idea and suggest a simplified and improved cerebellar model with priority-based delayed eligibility trace learning rule (S-CDE) that enables a mobile robot to freely and smoothly navigate in an environment. S-CDE is constructed in a brain-based device which mimics the anatomy, physiology, and dynamics of cerebellum. The input signal in terms of depth information generated from a simulated laser sensor is encoded as neuronal region activity for velocity and turn rate control. A priority-based delayed eligibility trace learning rule is proposed to maximize the usage of input signals for learning in synapses on Purkinje cell and cells in the deep cerebellar nuclei of cerebellum. Error signal generation and input signal conversion algorithms for turn rate and velocity are designed to facilitate training in an environment containing turns of varying curvatures. S-CDE is tested on a simulated mobile robot which had to randomly navigate maps of Singapore and Hong Kong expressways.
Inspired by infant development, we propose a three staged developmental framework for an anthropomorphic robot manipulator. In the first stage, the robot is initialized with a basic reach-and- enclose-on-contact movement capability, and discovers a set of behavior primitives by exploring its movement parameter space. In the next stage, the robot exercises the discovered behaviors on different objects, and learns the caused effects; effectively building a library of affordances and associated predictors. Finally, in the third stage, the learned structures and predictors are used to bootstrap complex imitation and action learning with the help of a cooperative tutor. The main contribution of this paper is the realization of an integrated developmental system where the structures emerging from the sensorimotor experience of an interacting real robot are used as the sole building blocks of the subsequent stages that generate increasingly more complex cognitive capabilities. The proposed framework includes a number of common features with infant sensorimotor development. Furthermore, the findings obtained from the self-exploration and motionese guided human-robot interaction experiments allow us to reason about the underlying mechanisms of simple-to-complex sensorimotor skill progression in human infants.
We model the autonomous development of brain-inspired circuits through two modalities-video stream and action stream that are synchronized in time. We assume that such multimodal streams are available to a baby through inborn reflexes, self-supervision, and caretaker's supervision, when the baby interacts with the real world. By autonomous development, we mean that not only that the internal (inside the "skull") self-organization is fully autonomous, but the developmental program (DP) that regulates the computation of the network is also task nonspecific. In this work, the task-nonspecificity is reflected by the fact that the actions associated with an attended object in a cluttered, natural, and dynamic scene is taught after the DP is finished and the "life" has begun. The actions correspond to neuronal firing patterns representing object type, object location and object scale, but learning is directly from unsegmented cluttered scenes. Along the line of where-what networks (WWN), this is the first one that explicitly models multiple "brain" areas-each for a different range of object scales. Among experiments, large natural video experiments were conducted. To show the power of automatic attention in unknown cluttered backgrounds, the last experimental group demonstrated disjoint tests in the presence of large within-class variations (object 3-D-rotations in very different unknown backgrounds), but small between-class variations (small object patches in large similar and different unknown backgrounds), in contrast with global classification tests such as ImageNet and Atari Games.
Discriminating between bipolar disorder (BD) and major depressive disorder (MDD) is a major clinical challenge due to the absence of known biomarkers; hence a better understanding of their pathophysiology and brain alterations is urgently needed. Given the complexity, feature selection is especially important in neuroimaging applications, however, feature dimension and model understanding present serious challenges. In this study, a novel feature selection approach based on linear support vector machine with a forward-backward search strategy (SVM-FoBa) was developed and applied to structural and resting-state functional magnetic resonance imaging data collected from 21 BD, 25 MDD and 23 healthy controls. Discriminative features were drawn from both data modalities, with which the classification of BD and MDD achieved an accuracy of 92.1% (1000 bootstrap resamples). Weight analysis of the selected features further revealed that the inferior frontal gyrus may characterize a central role in BD-MDD differentiation, in addition to the default mode network and the cerebellum. A modality-wise comparison also suggested that functional information outweighs anatomical by a large margin when classifying the two clinical disorders. This work validated the advantages of multimodal joint analysis and the effectiveness of SVM-FoBa, which has potential for use in identifying possible biomarkers for several mental disorders.
Cognitive workload is an important indicator of mental activity that has implications for human-computer interaction, biomedical and task analysis applications. Previously, subjective rating (self-assessment) has often been a preferred measure, due to its ease of use and relative sensitivity to the cognitive load variations. However, it can only be feasibly measured in a post-hoc manner with the user's cooperation, and is not available as an online, continuous measurement during the progress of the cognitive task. In this paper, we used a cognitive task inducing seven different levels of workload to investigate workload discrimination using electroencephalography (EEG) signals. The entropy, energy, and standard deviation of the wavelet coefficients extracted from the segmented EEGs were found to change very consistently in accordance with the induced load, yielding strong significance in statistical tests of ranking accuracy. High accuracy for subject-independent multichannel classification among seven load levels was achieved, across the twelve subjects studied. We compare these results with alternative measures such as performance, subjective ratings, and reaction time (response time) of the subjects and compare their reliability with the EEG-based method introduced. We also investigate test/re-test reliability of the recorded EEG signals to evaluate their stability over time. These findings bring the use of passive brain-computer interfaces (BCI) for continuous memory load measurement closer to reality, and suggest EEG as the preferred measure of working memory load.
Machine learning algorithms allow us to directly predict brain states based on functional magnetic resonance imaging (fMRI) data. In this study, we demonstrate the application of this framework to neuromarketing by predicting purchase decisions from spatio-temporal fMRI data. A sample of 24 subjects were shown product images and asked to make decisions of whether to buy them or not while undergoing fMRI scanning. Eight brain regions which were significantly activated during decision-making were identified using a general linear model. Time series were extracted from these regions and input into a recursive cluster elimination based support vector machine (RCE-SVM) for predicting purchase decisions. This method iteratively eliminates features which are unimportant until only the most discriminative features giving maximum accuracy are obtained. We were able to predict purchase decisions with 71% accuracy, which is higher than previously reported. In addition, we found that the most discriminative features were in signals from medial and superior frontal cortices. Therefore, this approach provides a reliable framework for using fMRI data to predict purchase-related decision-making as well as infer its neural correlates.
To investigate critical frequency bands and channels, this paper introduces deep belief networks (DBNs) to constructing EEG-based emotion recognition models for three emotions: positive, neutral and negative. We develop an EEG dataset acquired from 15 subjects. Each subject performs the experiments twice at the interval of a few days. DBNs are trained with differential entropy features extracted from multichannel EEG data. We examine the weights of the trained DBNs and investigate the critical frequency bands and channels. Four different profiles of 4, 6, 9 and 12 channels are selected. The recognition accuracies of these four profiles are relatively stable with the best accuracy of 86.65%, which is even better than that of the original 62 channels. The critical frequency bands and channels determined by using the weights of trained DBNs are consistent with the existing observations. In addition, our experiment results show that neural signatures associated with different emotions do exist and they share commonality across sessions and individuals. We compare the performance of deep models with shallow models. The average accuracies of DBN, SVM, LR and KNN are 86.08%, 83.99%, 82.70% and 72.60%, respectively.
Current EEG-based brain-computer interface technologies mainly focus on how to independently use SSVEP, motor imagery, P300, or other signals to recognize human intention and generate several control commands. SSVEP and P300 require external stimulus, while motor imagery does not require it. However, the generated control commands of these methods are limited and cannot control a robot to provide satisfactory service to the user. Taking advantage of both SSVEP and motor imagery, this paper aims to design a hybrid BCI system that can provide multimodal BCI control commands to the robot. In this hybrid BCI system, three SSVEP signals are used to control the robot to move forward, turn left, and turn right; one motor imagery signal is used to control the robot to execute the grasp motion. In order to enhance the performance of the hybrid BCI system, a visual servo module is also developed to control the robot to execute the grasp task. The effect of the entire system is verified in a simulation platform and a real humanoid robot, respectively. The experimental results show that all of the subjects were able to successfully use this hybrid BCI system with relative ease.
This paper focuses on electroencephalogram (EEG) manifestations of mental states and actions, emulation of control and communication structures using EEG manifestations, and their application in brain-robot interactions. The paper introduces a mentally emulated demultiplexer, a device which uses mental actions to demultiplex a single EEG channel into multiple digital commands. The presented device is applicable in controlling several objects through a single EEG channel. The experimental proof of the concept is given by an obstacle-containing trajectory which should be negotiated by a robotic arm with two degrees of freedom, controlled by mental states of a human brain using a single EEG channel. The work is presented in the framework of Human-Robot interaction (HRI), specifically in the framework of brain-robot interaction (BRI). This work is a continuation of a previous work on developing mentally emulated digital devices, such as a mental action switch, and a mental states flip-flop.
An agent tasked with solving a number of different decision making problems in similar environments has an opportunity to learn over a longer timescale than each individual task. Through examining solutions to different tasks, it can uncover behavioral invariances in the domain, by identifying actions to be prioritized in local contexts, invariant to task details. This information has the effect of greatly increasing the speed of solving new problems. We formalise this notion as action priors, defined as distributions over the action space, conditioned on environment state, and show how these can be learnt from a set of value functions. We apply action priors in the setting of reinforcement learning, to bias action selection during exploration. Aggressive use of action priors performs context based pruning of the available actions, thus reducing the complexity of lookahead during search. We additionally define action priors over observation features, rather than states, which provides further flexibility and generalizability, with the additional benefit of enabling feature selection. Action priors are demonstrated in experiments in a simulated factory environment and a large random graph domain, and show significant speed ups in learning new tasks. Furthermore, we argue that this mechanism is cognitively plausible, and is compatible with findings from cognitive psychology.
We developed a novel algorithm to estimate bias fields from brain magnetic resonance (MR) images using a gradient-based method. The bias field is modeled as a multiplicative and slowly varying surface. We fit the bias field by a low-order polynomial. The polynomial’s parameters are directly obtained by minimizing the sum of square errors between the gradients of MR images (both in the x-direction and y-direction) and the partial derivatives of the desired polynomial in the log domain. Compared to the existing retrospective algorithms, our algorithm combines the estimation of the gradient of the bias field and the reintegration of the obtained gradient polynomial together so that it is more robust against noise and can achieve better performance, which are demonstrated through experiments with both real and simulated brain MR images.
It is now widely accepted that concepts and conceptualization are key elements towards achieving cognition on a humanoid robot. An important problem on this path is the grounded representation of individual concepts and the relationships between them. In this article, we propose a probabilistic method based on Markov Random Fields to model a concept web on a humanoid robot where individual concepts and the relations between them are captured. In this web, each individual concept is represented using a prototype-based conceptualization method that we proposed in our earlier work. Relations between concepts are linked to the cooccurrences of concepts in interactions. By conveying input from perception, action, and language, the concept web forms rich, structured, grounded information about objects, their affordances, words, etc. We demonstrate that, given an interaction, a word, or the perceptual information from an object, the corresponding concepts in the web are activated, much the same way as they are in humans. Moreover, we show that the robot can use these activations in its concept web for several tasks to disambiguate its understanding of the scene.
Vision gives primates a wealth of information useful to manipulate the environment, but at the same time it can easily overwhelm their computational resources. Active vision is a key solution found by nature to solve this problem: a limited fovea actively displaced in space to collect only relevant information. Here we highlight that in ecological conditions this solution encounters four problems: 1) the agent needs to learn where to look based on its goals; 2) manipulation causes learning feedback in areas of space possibly outside the attention focus; 3) good visual actions are needed to guide manipulation actions, but only these can generate learning feedback; and 4) a limited fovea causes aliasing problems. We then propose a computational architecture (“BITPIC”) to overcome the four problems, integrating four bioinspired key ingredients: 1) reinforcement-learning fovea-based top-down attention; 2) a strong vision-manipulation coupling; 3) bottom-up periphery-based attention; and 4) a novel action-oriented memory. The system is tested with a simple simulated camera-arm robot solving a class of search-and-reach tasks involving color-blob “objects.” The results show that the architecture solves the problems, and hence the tasks, very efficiently, and highlight how the architecture principles can contribute to a full exploitation of the advantages of active vision in ecological conditions.
Exploring the functional mechanism of the human brain during semantics categorization and subsequently leverage current semantics-oriented multimedia analysis by functional brain imaging have been receiving great attention in recent years. In the field, most of existing studies utilized strictly controlled laboratory paradigms as experimental settings in brain imaging data acquisition. They also face the critical problem of modeling functional brain response from acquired brain imaging data. In this paper, we present a brain decoding study based on sparse multinomial logistic regression (SMLR) algorithm to explore the brain regions and functional interactions during semantics categorization. The setups of our study are two folds. First, we use naturalistic video streams as stimuli in functional magnetic resonance imaging (fMRI) to simulate the complex environment for semantics perception that the human brain has to process in real life. Second, we model brain responses to semantics categorization as functional interactions among large-scale brain networks. Our experimental results show that semantics categorization can be accurately predicted by both intrasubject and intersubject brain decoding models. The brain responses identified by the decoding model reveal that a wide range of brain regions and functional interactions are recruited during semantics categorization. Especially, the working memory system exhibits significant contributions. Other substantially involved brain systems include emotion, attention, vision and language systems.
Motion-onset visual evoked potential (mVEP) has been recently proposed for EEG-based brain-computer interface (BCI) system. It is a scalp potential of visual motion response, and typically composed of three components: P1, N2, and P2. Usually several repetitions are needed to increase the signal-to-noise ratio (SNR) of mVEP, but more repetitions will cost more time thus lower the efficiency. Considering the fluctuation of subject's state across time, the adaptive repetitions based on the subject's real-time signal quality is important for increasing the communication efficiency of mVEP-based BCI. In this paper, the amplitudes of the three components of mVEP are proposed to build a dynamic stopping criteria according to the practical information transfer rate (PITR) from the training data. During online test, the repeated stimulus stopped once the predefined threshold was exceeded by the real-time signals and then another circle of stimulus newly began. Evaluation tests showed that the proposed dynamic stopping strategy could significantly improve the communication efficiency of mVEP-based BCI that the average PITR increases from 14.5 bit/min of the traditional fixed repetition method to 20.8 bit/min. The improvement has great value in real-life BCI applications because the communication efficiency is very important.