Functional near-infrared spectroscopy (fNIRS) is a brain imaging technique used to estimate neuronal activity by measuring blood oxygenation. In this paper, we develop and evaluate an extensive set of fNIRS features for workload estimation, combining them with respiration and heartbeat signals. Our subject- and session-independent workload estimator is validated in a virtual flight simulator, where workload is objectively assessed based on task performance. We experiment with various regression models and feature ablations, identifying the most effective fNIRS features. The best fNIRS-based model achieves a correlation of 0.3188 with objective workload labels, improving to 0.3268 when incorporating breathing signals. This study demonstrates the value of our novel fNIRS feature set for workload estimation.
In this paper we present a workload estimator based on biological signals - electroencephalographic and eye-tracking. The workload estimator is person- and session-independent, designed to work in a virtual reality flight simulator environment and is a part of our adaptive training system. The novel component is using objective evaluation of the workload, based on the flight logs, as labels for training the regression neural network. As evaluation parameter is selected the correlation with the objective labels. The paper contains the results from using several feature sets and estimators, where the best estimator achieves correlation with objective labels of 0.84.
Silent speech recognition has emerged as a promising approach for enabling hands-free and discreet interaction with head-worn devices. In this paper, we present QuietSync, a multimodal system that combines inertial measurement unit (IMU) and contact electrode (ExG) signals to achieve accurate silent speech recognition using off-the-shelf devices. QuietSync utilizes an IMU attached to the lower part of the headphones near the ear and strategically places ExG electrodes on the headphones, glasses (nose and behind the ear), and face (for VR applications) to capture subtle movements and muscle activity associated with silent speech production. We conducted a user study with 9 participants and successfully recognized 12 commands with an accuracy of 94.2%. Our system leverages the complementary nature of IMU and ExG signals to enhance the robustness and reliability of silent speech recognition. The IMU captures subtle movements of the jaw and facial muscles, while the ExG electrodes detect low-amplitude surface muscle activity associated with speech production. We show that our system is not affected by the length and speech mannerisms of the commands, and can be fine-tuned for users of varied native languages with only 5 samples. Our findings demonstrate the feasibility of using off-the-shelf head-worn devices to enable silent speech recognition, opening up new possibilities for seamless and discreet interaction with devices such as VR/AR headsets and earables. To the best of our knowledge, QuietSync is the first system to enable silent speech interaction for multiple form factors.
Visual imagery, or the mental simulation of visual information from memory, could serve as an effective control paradigm for a brain-computer interface (BCI) due to its ability to directly convey the user’s intention with many natural ways of envisioning an intended action. However, multiple initial investigations into using visual imagery as a BCI control strategies have been unable to fully evaluate the capabilities of true spontaneous visual mental imagery. One major limitation in these prior works is that the target image is typically displayed immediately preceding the imagery period. This paradigm does not capture spontaneous mental imagery as would be necessary in an actual BCI application but something more akin to short-term retention in visual working memory. Results from the present study show that short-term visual imagery following the presentation of a specific target image provides a stronger, more easily classifiable neural signature in EEG than spontaneous visual imagery from long-term memory following an auditory cue for the image. We also show that short-term visual imagery and visual perception share commonalities in the most predictive electrodes and spectral features. However, visual imagery received greater influence from frontal electrodes whereas perception was mostly confined to occipital electrodes. This suggests that visual perception is primarily driven by sensory information whereas visual imagery has greater contributions from areas associated with memory and attention. This work provides the first direct comparison of short-term and long-term visual imagery tasks and provides greater insight into the feasibility of using visual imagery as a BCI control strategy.
In this paper, we propose a novel method to calculate trainee performance scores in a flight simulator environment using flight simulator logs. Our approach improves upon the existing scoring system designed in the AFRL by better fitting scores into the existing model of the training process.
Gaze tracking allows hands-free and voice-free interaction with computers, and has gained more use recently in virtual and augmented reality headsets. However, it traditionally uses dwell time for selection tasks, which suffers from the Midas Touch problem. Tongue gestures are subtle, accessible and can be sensed non-intrusively using an IMU at the back of the ear, PPG and EEG. We demonstrate a novel interaction method combining gaze tracking with tongue gestures for gaze-based selection faster than dwell time and multiple selection options. We showcase its usage as a point-and-click interface in three hands-free games and a musical instrument.
SSVEP-based BCIs are amongst the most promising BCIs in terms of speed and accuracy. However, despite significant effort from the community in order to make them more practical and user friendly, they remain particularly annoying to use. In this paper, we investigate the effect of the size and contrast of the SSVEP visual stimulations on both of the classification accuracy and the annoyance of the interface, with the global aim to find a trade-off between performance and user-friendliness. We conducted a user study on twelve (12) participants in order to evaluate the joint effect of different stimulation sizes and contrasts on the SSVEP classification accuracy in a Virtual Reality context. The results of this experiment suggest that the size of the stimulation has a significant impact on both of the classification accuracy, below a certain threshold, and on the perceived annoyance. No effect of the contrast was however found neither on the classification accuracy nor on the perceived annoyance, suggesting that it is still possible to accurately operate SSVEP-based BCIs using lower contrast stimulation.
Mouth-based interfaces are a promising new approach enabling silent, hands-free and eyes-free interaction with wearable devices. However, interfaces sensing mouth movements are traditionally custom-designed and placed near or within the mouth. TongueTap synchronizes multimodal EEG, PPG, IMU, eye tracking and head tracking data from two commercial headsets to facilitate tongue gesture recognition using only off-the-shelf devices on the upper face. We classified eight closed-mouth tongue gestures with 94% accuracy, offering an invisible and inaudible method for discreet control of head-worn devices. Moreover, we found that the IMU alone differentiates eight gestures with 80% accuracy and a subset of four gestures with 92% accuracy. We built a dataset of 48,000 gesture trials across 16 participants, allowing TongueTap to perform user-independent classification. Our findings suggest tongue gestures can be a viable interaction technique for VR/AR headsets and earables without requiring novel hardware.
Brain-computer interfaces (BCIs) employ various paradigms which afford intuitive, augmented control for users to navigate digital technologies. In this study we explore the application of these BCI concepts to predictive text systems: commonplace interactive and assistive tools with variable usage contexts and user behaviors. We conducted an experiment to analyze user neurophysiological responses under these different usage scenarios and evaluate the feasibility of a closed-loop, adaptive BCI for use with such technologies. We recorded electroencephalogram (EEG) and eye tracking (ET) data from participants while they completed a self-paced typing task in a simulated predictive text environment. Participants completed the task with different degrees of reliance on the predictive text system (completely dependent, completely independent, or their choice) and encountered both correct and incorrect text generations. Data suggest that erroneous text generations may evoke neurophysiological responses that can be measured with both EEG and pupillometry. Moreover, these responses appear to change according to users’ reliance on the predictive text system. Results show promise for use in a passive, hybrid, BCI with a closed-loop, adaptive framework, and support a neurophysiological approach to the challenge of real-time human feedback on system performance.
In this paper, we propose an adaptive training algorithm that accelerates the training process based on a parametric model of trainees and training scenarios. The proposed approach makes trial-by-trial recommendations on optimal scenario difficulty selections to maximize improvements in the trainee's absolute skill level.
In the last decade, the signal processing (SP) community has witnessed a paradigm shift from model-based to data-driven methods. Machine learning (ML)—more specifically, deep learning—methodologies are nowadays widely used in all SP fields, e.g., audio, speech, image, video, multimedia, and multimodal/multisensor processing, to name a few. Many data-driven methods also incorporate domain knowledge to improve problem modeling, especially when computational burden, training data scarceness, and memory size are important constraints.
Brain-Computer Interface (BCI) technology may provide individuals with motor impairments or even the general population a new way to interact with the world around them. However, current BCI systems using electroencephalography (EEG) can be unreliable and produce large variations in performance. Most studies seek to improve performance by focusing on signal processing and classification techniques. However, it may also be beneficial to investigate different control strategies. For this reason, the main objective of this pilot study was to investigate the use of visual imagery, a control paradigm that has not been much tested for EEG BCI applications. Visual imagery may provide a more intuitive control strategy with a greater number of available classes than other popular imagery-based methods such as motor imagery. Using this paradigm, we have demonstrated above chance binary classification accuracy (59.9%, p < 0.05) during offline decoding of face and scene visual imagery. Furthermore, the participant in this study achieved significantly above chance performance during a three-class, closed-loop BCI interaction (47.2%, p = 0.05). The initial results of this pilot study demonstrate the feasibility of using visual imagery as an alternative EEG BCI control paradigm.
In this paper we propose a parametric model of the training process which includes the skill increase and retention during training. The model is verified with the data from 22 subjects performing tasks in a flight simulator. The resulting computational model estimates the behavioral score.
Head worn displays are often used in situations where users’ hands may be occupied or otherwise unusable due to permanent or situational movement impairments. Hands-free interaction methods like voice recognition and gaze tracking allow accessible interaction with reduced limitations for user ability and environment. Tongue gestures offer an alternative method of private, hands-free and accessible interaction. However, past tongue gesture interfaces come in intrusive or otherwise inconvenient form factors preventing their implementation in head worn displays. We present a multimodal tongue gesture interface using existing commercial headsets and sensors only located in the upper face. We consider design factors for choosing robust and usable tongue gestures, introduce eight gestures based on the criteria and discuss early work towards tongue gesture recognition with the system.
Brain-computer interfaces (BCIs) using Electroencephalography (EEG) have drawn attention to providing alternative control pathways for users with motor disabilities or even the general public in real-world environments due to their robustness, relatively low cost, and high portability. However, EEG still suffers from large variability between subjects or between sessions of an individual subject. To obtain optimal performance, a BCI usually requires a user to go through a calibration process to fine-tune the model. This calibration process is usually long and could hinder the practicality of a BCI. In this study, we propose a closed-loop framework that monitors the user EEG responses to the action of a BCI. If an Error-related Potential (ErrP) is detected in the response, it is indicated that the BCI is making a wrong prediction. By using the information from this ErrP detector, we can include online testing trials into the training pool and further fine-tune the model over the time the BCI is used. Results suggest that the proposed framework can reach better results with a few additional trials when compared to the model pre-trained from some existing data. Also, the performance of the proposed model can gradually converge to a fully calibrated model, which suggests that the conventional calibration process could be replaced by online training.
Convolutional beamformers integrate the multichannel linear prediction model into beamformers, which provide good performance and optimality for joint dereverberation and noise reduction tasks. While longer filters are required to model long reverberation times, the computational burden of current online solutions grows fast with the filter length and number of microphones. In this work, we propose a low complexity convolutional beamformer using a Kalman filter derived affine projection algorithm to solve the adaptive filtering problem. The proposed solution is several orders of magnitude less complex than comparable existing solutions while slightly outperforming them on the REVERB challenge dataset.
In this talk we will make an overview of the acoustical design of the sound capture systems and discuss the general architecture of speech enhancement pipelines for the needs of distant speech recognition. The talk will discuss both classical algorithms using statistical signal processing and deep learning using neural networks. It will be illustrated with real-life examples from the acoustical design and speech enhancement audio pipelines in Kinect, HoloLens, and Microsoft Teams.
People enjoy listening to music as part of their life. This makes music an excellent choice for designing a user-friendly brain-computer interface (BCI) for long-term use. We propose a novel BCI system using music stimuli that relies on brain signals collected via Smartfones, an EEG recording device integrated into a pair of headphones. In a user study of the proposed system, participants were asked to pay attention to one of three musical instruments playing simultaneously from separate spatial directions. We used a stimulus reconstruction method to decode attention from EEG signals. Results show that the proposed system can achieve good decoding accuracy (>70%) while providing superior user-friendliness compared to a traditional EEG setup.