Linear models are becoming increasingly popular to investigate brain activity in response to continuous and naturalistic stimuli. In the context of auditory perception, these predictive models can be 'encoding', when stimulus features are used to reconstruct brain activity, or 'decoding' when neural features are used to reconstruct the audio stimuli. These linear models are a central component of some brain-computer interfaces that can be integrated into hearing assistive devices (e.g., hearing aids). Such advanced neurotechnologies have been widely investigated when listening to speech stimuli but rarely when listening to music. Recent attempts at neural tracking of music show that the reconstruction performances are reduced compared with speech decoding. The present study investigates the performance of stimuli reconstruction and electroencephalogram prediction (decoding and encoding models) based on the cortical entrainment of temporal variations of the audio stimuli for both music and speech listening. Three hypotheses that may explain differences between speech and music stimuli reconstruction were tested to assess the importance of the speech-specific acoustic and linguistic factors. While the results obtained with encoding models suggest different underlying cortical processing between speech and music listening, no differences were found in terms of reconstruction of the stimuli or the cortical data. The results suggest that envelope-based linear modelling can be used to study both speech and music listening, despite the differences in the underlying cortical mechanisms.
It has been demonstrated that from cortical recordings, it is possible to detect which speaker a person is attending in a cocktail party scenario. The stimulus reconstruction approach, based on linear regression, has been shown to be useable to reconstruct an approximation of the envelopes of the sounds attended to and not attended to by a listener from the electroencephalogram data (EEG). Comparing the reconstructed envelopes with the envelopes of the stimuli, a higher correlation between the envelopes of the attended sound is observed. Most of the studies focused on speech listening, and only a few studies investigated the performances and the mechanisms of auditory attention decoding during music listening. In the present study, auditory attention detection (AAD) techniques that have been proven successful for speech listening were applied to a situation where the listener is actively listening to music concomitant with a distracting sound. Results show that AAD can be successful for both speech and music listening while showing differences in the reconstruction accuracy. The results of this study also highlighted the importance of the training data used in the construction of the model. This study is a first attempt to decode auditory attention from EEG data in situations where music and speech are present. The results of this study indicate that linear regression can also be used for AAD when listening to music if the model is trained for musical signals.
This dataset contains EEG recordings from 18 subjects listening to continuous sound, either speech or music. Continuous audio stimuli were presented to listeners in trials of 70 seconds from one loudspeaker located 150 cm in from of them. They were instructed to attentively listen to the sound during the whole trial. All listeners were native Danish speakers and were presented with 5 different types of audio stimuli: Instrumental music: Excerpt of Disney songs: Excerpts of polyphonic Disney songs with no lyrics. The melody line from the original version was replaced by a similar melody played by a synthetic cello (Referred to as MC: Music Cello in the dataset). Music with understood lyrics: Excerpts of polyphonic Disney songs with lyrics in Danish, understood by the listeners (Referred to as MD: Music Danish in the dataset). Music with non-understood lyrics: Excerpts of polyphonic Disney songs with lyrics in Finnish, not understood by the listeners (Referred to as MF: Music Finnish in the dataset). Understood speech: Excerpts of an audiobook in Danish read by a woman, understood by the listeners (Referred to as SD: Speech Danish in the dataset). Non-Understood speech: Excerpts of an audiobook in Finnish read by a woman, understood by the listeners (Referred to as SF: Speech Danish in the dataset). Data were recorded using a 64-channels g.HIamp-Research system and digitalized at a sampling rate of 2400 Hz. The dataset contains pre-processed EEG data (see pre-processing step applies to the data below), for each listener. Trials with large noise artefacts have been removed. The processed folder contains data used in Simon, A. et al. (2022) Cortical linear encoding and decoding of sounds: Differences between naturalistic speech and music listening. (Submitted). The processed data contains EEG data and an aligned audio envelope for each category of audio stimuli. The MATLAB script contains the processing applied to obtain it. The dataset was created within the InHear project. For more information, amds@es.aau.dk Preprocessing done -re-reference to average channels -downsampling to 512Hz -bandpass filter 0.5-45 Hz -ICA decomposition using SOBI algorithm -removed eyes and noise components
In complex sound scenes, where multiple sounds are present around a listener, selective attention to one auditory stream is hypothesized to synchronize low-frequency brain activity with the envelope of the attended streams. Recent research has employed stimulus reconstruction from neural data to decode to which auditory stream a listener is paying attention. This could be used to create an auditory attention decoder (AAD), that could be embedded in smart headphones or hearing aids, that would adapt the sound processing based on the attention of the user. However, most of these studies use full scalp electroencephalogram, which is not suitable for implementations in audio devices. To that aim, a smaller EEG device, with fewer electrodes could be used. In the present study, we explore the performance of an AAD based on a smaller number of electrodes during speech and music listening. Participants were presented with two sounds simultaneously, and where asked to attend to one while ignoring the other, and their cortical response was continuously recorded during the lsitening. Using a greedy approach based on reconstruction accuracy, a subset of EEG electrodes that are optimized for linear stimulus reconstruction were selected. The goal of this study is to explore the performance of a linear AAD when reducing the number of electrodes. Results suggest that four well-selected electrodes can be sufficient for a miniaturized AAD as it performs as well as a 64-channels setup. The channels selected vary depending on the type of sound attended, suggesting that different electrodes placement should be used to decode attention during music listening and speech listening.
The use of the term immersion to describe a multitude of varying experiences in the absence of a definitional consensus has obfuscated and diluted the term. The non-exhaustive literature review presented in this paper indicates that immersion is a psychological concept as opposed to being a property of the system or technology that facilitates an experience. An adaptable definition of immersion is synthesized based on the findings from the literature review: a state of deep mental involvement in which the individual may experience disassociation from the awareness of the physical world due to a shift in their attentional state. This definition is used to contrast and differentiate interchangeably used terms such as presence from immersion and outline the implications for conducting immersion research on audiovisual experiences. A new methodology for quantifying immersion is proposed and avenues for future work are briefly discussed.
We carry out a study to investigate how naïve (non-audio-experts) users understand the concept of sound directivity. This was done in the context of loudspeaker reproduction of sound fields with the purpose being to discover how such acoustical phenomena can best be explained to users. We investigated the mental models of 20 participants via an interview-based approach, in which we asked participants to draw and explain how they understood directivity, only providing them with minimal prior information. The interviews and drawings were analysed and mental models were extracted. Our analysis showed the models could be categorized into three General Mental Model Types (GMMTs): Direction of Sound, Area with Sound, and Sound Waves. These GMMTs were then used to build a 3-level combined model that also contains observations of what each GMMT is suitable for when trying to e.g. explain or illustrate loudspeaker directivity to non-expert users. These guidelines can be useful for designing visual representations to help explain loudspeaker directivity to non-expert users, and could be used for visualising further complex acoustical concepts.