Limited training data constrains deep learning models for Auditory Attention Decoding (AAD) in hearing aids (HAs). AAD uses electroencephalogram (EEG) data to decode listener's attention, enabling real-time tracking of specific sound sources. However, achieving high AAD performance with short time windows typical in HAs (<=1s) is challenging due to the scarcity of real-world speech-evoked EEG data. To address this issue, we investigate diffusion probabilistic models (DPMs) for generating synthetic speech-evoked EEG data. DPMs learn the underlying complex data structure through a denoising process and can generate realistic samples suitable for data augmentation. We evaluate the use of synthetic EEG data for augmenting datasets in locus-of-attention (LoA) classification tasks. Our experiments demonstrate that DPMs can generate realistic EEG signals and that incorporating synthetic data significantly improves AAD performance compared to models trained solely on measured EEG data (p<0.05). These results highlight the potential of diffusion-based data augmentation to mitigate training data limitations and improve the robustness of short-window AAD models in HA applications.
Carefully selecting the source data is crucial to achieve high performance of transfer learning methods for brain-computer interfaces (BCIs). Especially so in settings where a large amount of source data is available, and finding the optimal source is not computationally feasible. This paper presents a novel method for source selection, the so-called Transfer Performance Predictor (TPP) method. The TPP method is based on computationally simple features, a choice made to enable real-time implementation and reduce calibration time. The presented method outperforms other comparable source selection methods in BCI settings where a large amount of source data is available. By using the TPP method, source selection can be performed quickly with good results for transfer learning performance, which means that the BCI calibration time can be reduced and a new target user can more quickly start using the BCI.
This paper introduces the pole ratio metric and presents a sphere-based view of symmetric positive-definite matrix rotations on the Riemannian manifold of symmetric positive-definite matrices equipped with the affine-invariant Riemannian metric. The pole ratio quantifies whether data from different users lie on this Riemannian manifold in a way that enables effective transfer learning. The sphere-based view provides insight into the rotational step of transfer learning using the Riemannian Procrustes analysis method and highlights the limitations of rotation. For effective transfer learning, selecting appropriate source data is essential for good performance. The pole ratio is shown to be an effective metric for selecting source data. The main contribution of the paper is the insight into the limitations of rotations on a Riemannian manifold; the usefulness of the pole ratio as a source selection metric is a natural extension of this insight. This paper focuses on Brain-Computer Interfaces (BCIs), but the sphere-based view of rotations of symmetric positive-definite matrix data and the pole ratio are applicable to any field that models two-class data using symmetric positive-definite matrices.
This study investigates the neural encoding of speech features in hearing aid users using electroencephalography (EEG) during a simulated cocktail party scenario. The objective was to investigate neural tracking of various acoustic and linguistic features and how hearing aid noise reduction influenced this tracking. The features analyzed included the acoustic envelope, phonetic features, word onset, and word surprisal, the latter derived from GPT-2. Temporal Response Functions (TRFs) were used to correlate these features with EEG signals, revealing how the brain tracks attended (target) versus unattended (masker) speech. TRFs were estimated using a boosting algorithm, with speech features as predictors and EEG signals as responses. Results revealed a significant distinction between target and masker speech. The acoustic envelope showed the strongest correlation with EEG responses. Distinct tracking patterns were observed: the acoustic envelope and phonetic features correlated with early processing stages, while word onset and word suprisal were linked to later stages. Noise reduction further influenced the tracking of these features. These findings improve our understanding of how hearing aid users process speech and provide insight for developing hearing aids that adapt to individual neural responses.
Artificial intelligence advances have recently influenced wireless communications, including beam management in fifth-generation (5G) new radio systems. AI-driven models and algorithms are being applied to enhance tasks such as beam selection, prediction, and refinement by leveraging real-time and historical data. These approaches address challenges such as mobility under complex channel conditions, showing promising results compared to traditional methods. Beam management in 5G refers to processes that ensure optimal alignment between the base station and user equipment for effective signal transmission and reception based on real-time channel state information and user positioning. This study leverages accurate beam prediction to identify a smaller subset of beams, resulting in a more efficient, streamlined, and link-adaptive communication system. The innovative approach presented introduces a precise, attention-based prediction model that derives the entire downlink transmission chain in a commercial grade 5G system. The predicted downlink beams are specifically tailored to handle the complexities of none line-of-sight environments known for high-dimensional channel dynamics and scatterer-induced signal variations. This novel method introduces a paradigm shift in utilizing environmental and channel dynamics in contrast to conventional procedures of beam management, which entails complex methods involving exhaustive techniques to predict the best beams. The presented beam prediction results demonstrate robustness in addressing the challenges posed by signal-dispersive environments, showcasing great potential in mobility scenarios.
Objective. This study aimed to investigate the potential of contrastive learning to improve auditory attention decoding (AAD) using electroencephalography (EEG) data in challenging cocktail-party scenarios with competing speech and background noise.Approach. Three different models were implemented for comparison: a baseline linear model (LM), a non-LM without contrastive learning (NLM), and a non-LM with contrastive learning (NLMwCL). The EEG data and speech envelopes were used to train these models. The NLMwCL model used SigLIP, a variant of CLIP loss, to embed the data. The speech envelopes were reconstructed from the models and compared with the attended and ignored speech envelopes to assess reconstruction accuracy, measured as the correlation between the reconstructed and actual speech envelopes. These reconstruction accuracies were then compared to classify attention. All models were evaluated in 34 listeners with hearing impairment.Results. The reconstruction accuracy for attended and ignored speech, along with attention classification accuracy, was calculated for each model across various time windows. The NLMwCL consistently outperformed the other models in both speech reconstruction and attention classification. For a 3-second time window, the NLMwCL model achieved a mean attended speech reconstruction accuracy of 0.105 and a mean attention classification accuracy of 68.0%, while the NLM model scored 0.096 and 64.4%, and the LM achieved 0.084 and 62.6%, respectively.Significance. These findings demonstrate the promise of contrastive learning in improving AAD and highlight the potential of EEG-based tools for clinical applications, and progress in hearing technology, particularly in the design of new neuro-steered signal processing algorithms.
We address the recognized person-to-person Brain-Computer Interface (BCI) calibration problem and tackle session-dependency through the use of unsupervised canonical polyadic (CP) tensor decomposition. For a motor imagery task, the approach reveals universal structures within EEG data, common between subjects and prominent for a certain task. Further, we develop a novel similarity measure that includes weighting of the decomposition's factor matrices, and argue that it is more representative than what has previously been presented in literature. The proposed similarity measure shows potential in a BCI classification task, i.e. drowsiness during simulated driving (average Pearson correlation of 0.6).
The integration of high-precision cellular localization and machine learning (ML) is considered a cornerstone technique in future cellular navigation systems, offering unparalleled accuracy and functionality. This study focuses on localization based on uplink channel measurements in a fifth-generation (5G) new radio (NR) system. An attention-aided ML-based single-snapshot localization pipeline is presented, which consists of several cascaded blocks, namely a signal processing block, an attention-aided block, and an uncertainty estimation block. Specifically, the signal processing block generates an impulse response beam matrix for all beams. The attention-aided block trains on the channel impulse responses using an attention-aided network, which captures the correlation between impulse responses for different beams. The uncertainty estimation block predicts the probability density function of the user equipment (UE) position, thereby also indicating the confidence level of the localization result. Two representative uncertainty estimation techniques, the negative log-likelihood and the regression-by-classification techniques, are applied and compared. Furthermore, for dynamic measurements with multiple snapshots available, we combine the proposed pipeline with a Kalman filter to enhance localization accuracy. To evaluate our approach, we extract channel impulse responses for different beams from a commercial base station. The outdoor measurement campaign covers Line-of-Sight (LoS), Non Line-of-Sight (NLoS), and a mix of LoS and NLoS scenarios. The results show that sub-meter localization accuracy can be achieved.
Objective.This study develops a deep learning (DL) method for fast auditory attention decoding (AAD) using electroencephalography (EEG) from listeners with hearing impairment (HI). It addresses three classification tasks: differentiating noise from speech-in-noise, classifying the direction of attended speech (left vs. right) and identifying the activation status of hearing aid noise reduction algorithms (OFF vs. ON). These tasks contribute to our understanding of how hearing technology influences auditory processing in the hearing-impaired population.Approach.Deep convolutional neural network (DCNN) models were designed for each task. Two training strategies were employed to clarify the impact of data splitting on AAD tasks: inter-trial, where the testing set used classification windows from trials that the training set had not seen, and intra-trial, where the testing set used unseen classification windows from trials where other segments were seen during training. The models were evaluated on EEG data from 31 participants with HI, listening to competing talkers amidst background noise.Main results.Using 1 s classification windows, DCNN models achieve accuracy (ACC) of 69.8%, 73.3% and 82.9% and area-under-curve (AUC) of 77.2%, 80.6% and 92.1% for the three tasks respectively on inter-trial strategy. In the intra-trial strategy, they achieved ACC of 87.9%, 80.1% and 97.5%, along with AUC of 94.6%, 89.1%, and 99.8%. Our DCNN models show good performance on short 1 s EEG samples, making them suitable for real-world applications. Conclusion: Our DCNN models successfully addressed three tasks with short 1 s EEG windows from participants with HI, showcasing their potential. While the inter-trial strategy demonstrated promise for assessing AAD, the intra-trial approach yielded inflated results, underscoring the important role of proper data splitting in EEG-based AAD tasks.Significance.Our findings showcase the promising potential of EEG-based tools for assessing auditory attention in clinical contexts and advancing hearing technology, while also promoting further exploration of alternative DL architectures and their potential constraints.
The performance of HVAC equipment, including chillers, is continuing to be pushed to theoretical limits, which impacts the necessity for advanced control logic to operate them efficiently and robustly. At the same time, their architectures are becoming more complex; many systems have multiple compressors, expansion devices, evaporators, circuits, or other elements that challenge control design and resulting performance. In order to maintain respectful controlled speed of response, stability, and robustness, controllers are becoming more complex, including the move from thermostatic control, to proportional integrator (PI), and to multiple-input multiple-output (MIMO) controllers. Model-based control design works well for their synthesis, while having accurate models for numerous product variants is unrealistic, often leading to very conservative designs. To address this, we propose and demonstrate a learning-based control tuner that supports the design of MIMO decoupling PI controllers using online information to adapt controller coefficients from an initial guess during commissioning or operation. The approach is tested on a physics-based model of a water-cooled screw chiller. The method is able to find a controller that performs better than a nominal controller (two single PI controllers) in terms of decreasing deviations from the operating point during disturbances while still following reference changes.
Objective.This paper presents a novel domain adaptation (DA) framework to enhance the accuracy of electroencephalography (EEG)-based auditory attention classification, specifically for classifying the direction (left or right) of attended speech. The framework aims to improve the performances for subjects with initially low classification accuracy, overcoming challenges posed by instrumental and human factors. Limited dataset size, variations in EEG data quality due to factors such as noise, electrode misplacement or subjects, and the need for generalization across different trials, conditions and subjects necessitate the use of DA methods. By leveraging DA methods, the framework can learn from one EEG dataset and adapt to another, potentially resulting in more reliable and robust classification models.Approach.This paper focuses on investigating a DA method, based on parallel transport, for addressing the auditory attention classification problem. The EEG data utilized in this study originates from an experiment where subjects were instructed to selectively attend to one of the two spatially separated voices presented simultaneously.Main results.Significant improvement in classification accuracy was observed when poor data from one subject was transported to the domain of good data from different subjects, as compared to the baseline. The mean classification accuracy for subjects with poor data increased from 45.84% to 67.92%. Specifically, the highest achieved classification accuracy from one subject reached 83.33%, a substantial increase from the baseline accuracy of 43.33%.Significance.The findings of our study demonstrate the improved classification performances achieved through the implementation of DA methods. This brings us a step closer to leveraging EEG in neuro-steered hearing devices.
Cellular user positioning is a promising service provided by Fifth Generation New Radio (5G NR) networks.Besides, Machine Learning (ML) techniques are foreseen to become an integrated part of 5G NR systems improving radio performance and reducing complexity.In this paper, we investigate ML techniques for positioning using 5G NR fingerprints consisting of uplink channel estimates from the physical layer channel.We show that it is possible to use Sounding Reference Signals (SRS) channel fingerprints to provide sufficient data to infer user position.Furthermore, we show that small fully-connected moderately Deep Neural Networks, even when applied to very sparse SRS data, can achieve successful outdoor user positioning with meter-level accuracy in a commercial 5G environment.
This paper introduces an improved method for real-time brain computer interface control. We demonstrate how Bayesian optimization and feedback can be used to achieve faster statistical convergence by controlling the sequence of stimuli shown in a brain computer interface based on a visual oddball paradigm.
In dual control, the manipulated variables are used to both regulate the system and identify unknown parameters. The joint probability distribution of the system state and the parameters is known as the hyperstate. The paper proposes a method to perform dual control using a deep reinforcement learning algorithm in combination with a neural network model trained to represent hyperstate transitions. The hyperstate is compactly represented as the parameters of a mixture model that is fitted to Monte Carlo samples of the hyperstate. The representation is used to train a hyperstate transition model, which is used by a standard reinforcement learning algorithm to find a dual control policy. The method is evaluated on a simple nonlinear system, which illustrates a situation where probing is needed, but it can also scale to high-dimensional systems. The method is demonstrated to be able to learn a probing technique that reduces the uncertainty of the hyperstate, resulting in improved control performance.
The multi-armed bandit (MAB) problem models a decision-maker that optimizes its actions based on current and acquired new knowledge to maximize its reward. This type of online decision is prominent in many procedures of Brain-Computer Interfaces (BCIs) and MAB has previously been used to investigate, e.g., what mental commands to use to optimize BCI performance. However, MAB optimization in the context of BCI is still relatively unexplored, even though it has the potential to improve BCI performance during both calibration and real-time implementation. Therefore, this review aims to further describe the fruitful area of MABs to the BCI community. The review includes a background on MAB problems and standard solution methods, and interpretations related to BCI systems. Moreover, it includes state-of-the-art concepts of MAB in BCI and suggestions for future research.
The matched phase reassignment, developed to estimate phase synchrony of transient oscillatory signals, is extended into a multitaper phase reassignment (MTPR) method. The method gives perfect time-frequency localization for two transients with zero phase difference and estimates of time locations and oscillatory frequencies in low signal-to-noise ratios. For different signal-to-noise ratios between channels a suggestion of corrected reassignment vector expressions is given, resulting in minimized variance. The MTPR outperforms the matched phase reassignment as well as state-of-the-art methods, such as Pearson's linear correlation, time-frequency cross-spectrogram phase estimation and the Phase Lag Index method. An example of estimated phase differences, time locations and oscillatory frequencies of electrical signals measured from the brain is also shown.
We demonstrate that finite impulse response (FIR) models can be applied to analyze the time evolution of an epidemic with its impact on deaths and healthcare strain. Using time series data for COVID-19-related cases, ICU admissions and deaths from Sweden, the FIR model gives a consistent epidemiological trajectory for a simple delta filter function. This results in a consistent scaling between the time series if appropriate time delays are applied and allows the reconstruction of cases for times before July 2020, when RT-PCR testing was not widely available. Combined with randomized RT-PCR study results, we utilize this approach to estimate the total number of infections in Sweden, and the corresponding infection-to-fatality ratio (IFR), infection-to-case ratio (ICR), and infection-to-ICU admission ratio (IIAR). Our values for IFR, ICR and IIAR are essentially constant over large parts of 2020 in contrast with claims of healthcare adaptation or mutated virus variants importantly affecting these ratios. We observe a diminished IFR in late summer 2020 as well as a strong decline during 2021, following the launch of a nation-wide vaccination program. The total number of infections during 2020 is estimated to 1.3 million, indicating that Sweden was far from herd immunity.