Background/Objectives: Adversarial domain adaptation methods are widely used in EEG-based emotion recognition to reduce the influence of individual differences and the non-stationary characteristics of electroencephalogram (EEG) signals. Most existing methods employ binary domain discriminators to align source and target domains at the global distribution level. However, such strategies often neglect the potential multimodal structure of emotional EEG data and the asymmetric emotional processing characteristics of the left and right hemispheres. To address these issues, this study proposes a Bi-Hemispheric Adversarial Domain Adaptation Neural Network (BiHADA) for EEG-based emotion recognition. Methods: In the proposed BiHADA framework, the conventional binary domain discriminator is extended into a multimodal discriminator by incorporating the label structure information of source-domain data into the domain discrimination process. This design encourages features belonging to the same emotional category to be aligned across domains and promotes positive knowledge transfer. In addition, dual adversarial domain adaptation branches are constructed to model the left and right hemispheres separately, enabling the network to capture hemisphere-specific emotional representations. Furthermore, discriminator-derived perplexity is introduced to evaluate the distribution alignment quality of target samples and to adaptively determine the weights of the corresponding hemisphere classifiers, thereby reducing the influence of poorly aligned samples during the final decision stage. Results: Experiments on the SEED dataset show that BiHADA achieves classification accuracies of 86.82% and 92.71% in cross-subject and cross-session tasks, respectively. These results demonstrate that the proposed method can effectively improve the transferability and discriminability of EEG emotional features under different domain adaptation scenarios. Conclusions: The proposed BiHADA method enhances EEG-based emotion recognition by jointly considering class-structure-guided domain alignment, hemispheric functional asymmetry, and branch-wise adaptation quality. The results suggest that incorporating source-domain label structure and hemisphere-specific adaptation can improve cross-domain EEG emotion recognition performance.
Graph convolutional neural networks (GCNs) have gained popularity in electroencephalogram (EEG) recognition research for their ability to capture topological features. However, deeper GCNs are prone to over-smoothing, which can significantly reduce the accuracy and effectiveness of EEG recognition. To address this problem, we propose a multilayer GCN model for motor imagery classification. This approach aims to effectively extract multilevel topological information while preserving node differentiation. The multilayer GCN architecture integrates information across levels through a concatenation module that merges the outputs of each GCN layer in both the time and frequency branches. Furthermore, the incorporation of a channel location encoder and a channel self-attention graph embedding module significantly improves the model’s ability to adaptively capture dynamic relationships between channels, enhancing the fusion of topological and spatio-temporal features. In addition, the adjacency matrix of the graph is dynamically updated via a learnable weight matrix, allowing for the adaptive learning of topological relationships among electrodes. Finally, evaluated using the BCI Competition IV 2a and OpenBMI datasets, the proposed model achieves classification accuracies of 85.1% and 70.0%, respectively, outperforming other GCN-based methods. Experimental results indicate that this GCN-based approach addresses over-smoothing and has the potential to improve BCI performance for motor rehabilitation.
Understanding emotional states is fundamental to advancing next-generation AI systems with human-like attributes. Combining the complementary strengths of electroencephalog raphy (EEG) and facial expressions holds great promise for advancing multimodal emotion recognition (MER). EEG provides objective measurements of neural activity, while facial expressions convey rich, externally observable emotional cues. However, existing joint learning frameworks often fall short of fully exploiting the synergy between these modalities. Two key challenges remain unresolved: (1) insufficient cross-modal alignment and interaction prior to fusion which limits the semantic complementarity between modalities; and (2) modality unreliability caused by temporal fluctuations in signal quality and inconsistencies in emotional semantics. To address these limitations, we propose a novel framework that integrates a Step-wise Prompts (SwiP) module with an Uncertainty-Aware Dynamic Fusion (UADF) mechanism. SwiP enables progressive, fine-grained interaction by introducing sequential facial features as visual prompts to guide EEG representation learning, thereby enhancing cross modal complementarity. UADF dynamically adjusts modality contributions through a token- and modality-level uncertainty estimation scheme, enabling the model to selectively emphasize informative inputs and suppress noisy or irrelevant signals. Extensive experiments on benchmark datasets demonstrate that our method achieves state-of-the-art performance, consistently outperforming competitive baselines in both accuracy and stability. These results highlight the potential of our framework as a robust and generalizable solution for real-world affective computing applications.
Emotions are constantly generated in daily activities. They not only control people's behavioral patterns and thinking decisions, but also affect physical and mental health. Therefore, the emotion recognition technology based on electroencephalogram (EEG) signals has broad application prospects in multiple fields such as human-computer interaction, medical health, and intelligent driving. To address the issues of insufficient labeled data in EEG and significant differences in EEG data among different subjects and at different time periods, this paper introduces the domain adaptation (DA) technique to solve the task of cross-domain EEG emotion recognition under unsupervised conditions. Aiming at the problem that the existing domain adaptation methods ignore the different weights of the global domain and subdomains when reducing domain differences, this paper proposes a dynamic bi-domain discriminator adversarial network (DBDAN). A feature extractor is constructed to extract the low-level domain invariant features of the EEG signals, and a multi-branch domain-specific feature extractor is used to generate domain-specific features. In order to reduce domain differences, adversarial learning for the global domain and subdomains is achieved through a dual-domain discriminator. Meanwhile, dynamic factors are introduced to dynamically adjust adversarial learning, enabling the model to strike a balance between coarse-grained adversarial and fine-grained adversarial and achieve domain adaptation more precisely. The cross-subject accuracy rates on SEED, SEED-IV and DEAP were 89.43%, 75.09% and 63.42%, respectively. These results indicate that our method achieves competitive performance among the evaluated EEG transfer-learning baselines and suggest that dynamic global-subdomain alignment is beneficial for cross-domain EEG emotion recognition.
Goal: Deep learning-based motor imagery EEG classification is limited by data scarcity, which constrains model generalization and performance. Methods: We propose a dual-cascade generative adversarial network (dcGAN) framework with a variable focused attention (VFA) module for MI-EEG data augmentation. The first stage learns latent frequency-domain priors from random noise through an adversarial training scheme; the second stage then synthesizes artificial EEG samples with a U-Net generator conditioned on these priors, augmented by the VFA module and a time-domain consistency loss. A VFA-enhanced EEGNet is subsequently trained on the combination of real and generated samples for classification. Results: On the BCI Competition IV 2a and 2b datasets, the proposed method achieves classification accuracies of 84.92% and 91.79%, with Cohen’s Kappa coefficients of 0.79 and 0.81, respectively, outperforming baseline methods. Conclusions: The integration of structured frequency-domain priors and attention mechanisms improves the fidelity of generated EEG samples, which in turn enhances downstream classification performance.
Background: Spiking neural networks (SNNs) have attracted significant attention in the field of brain-computer interfaces owing to their distinctive biological plausibility and energy efficiency advantages. However, the discrete nature of spikes renders gradient-based differentiation infeasible, making it difficult to directly obtain well-trained SNNs. A common approach is to transfer the weights from artificial neural networks (ANNs) to SNNs. However, this process introduces conversion errors that pose significant challenges. Methods: To address these challenges, we propose the self-rectifying integrate-and-fire (SRIF) neuron, which employs negative spikes to reduce asynchronism error and rectification spikes to diminish clipping error. Concomitantly, we propose a collaborative trim (CT) training framework that introduces a quantized network to perceive the weights and results of SNNs, which can further improve performance. Result: The proposed training methodology enables SNNs to achieve performance metrics comparable to those of ANNs in EEG-based motor imagery (MI) classification. Conclusions: Experimental results demonstrate that our method not only preserves the superior classification performance of ANNs but also leverages the superior energy efficiency and lower computational complexity of SNNs.
Accurate differentiation of Alzheimer’s disease (AD) and frontotemporal dementia (FTD) is clinically important because their management strategies differ. This study aimed to develop and evaluate a Dynamic Threshold Graph Convolutional Network (DT-GCN) and to systematically compare four EEG-based functional connectivity (FC) measures—Pearson correlation, phase-locking value (PLV), Granger causality, and copula analysis—for classifying AD, FTD, and healthy controls (HCs). In contrast to fixed graph binarization, DT-GCN updates the connectivity threshold at each training epoch according to training-fold loss. The model is trained under a multi-task objective combining classification, reconstruction, and contrastive losses, with the loss weights adjusted across three stages of training. Resting-state 19-channel EEG recordings from a public dataset (DS004504) comprising 36 AD, 23 FTD, and 29 HC participants were segmented into non-overlapping 8 s epochs. FC estimates were aggregated to obtain a subject-level adjacency matrix for each measure, with one-hot node identity and node degree as node features. Classification was performed at the subject level using stratified five-fold cross-validation with normalization and hyperparameter selection confined to training folds. Under the broadband setting, copula-based FC with DT-GCN yielded the highest observed three-class accuracy of 0.82±0.07 (macro-F1 0.79±0.07), compared with 0.45±0.06 accuracy and 0.36±0.06 macro-F1 for the standard GCN baseline. However, the small single-center cohort (N=88, including 23 participants with FTD) and the resulting small test folds limit the precision of the performance estimates and preclude robust inferential comparisons among FC methods and classification scenarios. These exploratory, dataset-specific findings require independent external validation and should not be interpreted as generalizable clinical performance or validated clinical biomarkers.
Retinal vessel segmentation is a fundamental task in automated retinal image analysis and plays a crucial role in assessing various ophthalmic diseases. However, accurate segmentation remains challenging due to extreme vessel scale variations, disrupted topological continuity, and the presence of noise and low-contrast regions in fundus images. To address these challenges, this paper proposes WMKA-Net, a Weighted Multi-Kernel Attention Network with a dual-stage architecture that integrates the Residual Multi-Scale Fusion Module (RMS) and Vascular-Oriented Attention Module (VOAM).The RMS module aggregates multi-scale vascular features through depth-specific kernel configurations and fusion strategies, with branch complexity statically designed to match hierarchical vessel calibers (from capillaries to main vessels). The VOAM module is specifically designed for retinal vessel segmentation and employs a dual-path attention mechanism to model long-range axial continuity and enhance bifurcation regions, thereby mitigating vessel fragmentation and improving structural completeness.Extensive experiments conducted on five public retinal vessel segmentation benchmarks (DRIVE, STARE, CHASE_DB1, HRF, and IOSTAR) demonstrate that WMKA-Net achieves competitive and well-balanced performance across multiple evaluation metrics, including sensitivity, specificity, F1-score, accuracy, and Area Under the Curve (AUC). These results indicate that the proposed method provides an effective and robust framework for retinal vessel segmentation and may serve as a technical foundation for automated retinal vessel analysis in clinical screening workflows.
Individual differences and non-stationary characteristics are prominent in EEG signals. Therefore, aligning the source and target domain data becomes essential in cross-subject and cross-session classification tasks. Although many adversarial adaptation networks can achieve distribution alignment through domain-level adaptation, they tend to disregard the multi-modal structure inherent in the data. In this paper, we present a concise and effective adversarial paradigm for EEG emotion recognition. This approach fully utilizes the label structure information of source domain data to reuse the binary discriminator as a class-informed discriminator instead of introducing additional modules, which not only realizes domain confusion, but also ensures that mode information is retained in the process of confusion to avoid mode collapse. To evaluate our method, a systematic experimental study was conducted on the public datasets SEED and SEED-IV. The average accuracy of cross-subject and cross-session scenarios achieved 90.21%, 95.47% on SEED, and 77.50%, 77.54% on SEED-IV respectively. Compared to the existing domain adaptation methods, the evident improvements of classification performance demonstrate the feasibility and effectiveness of our method.
The classification of attention states utilizing both electroencephalography (EEG) and functional near-infrared spectroscopy (fNIRS) is pivotal in understanding human cognitive functions. While multimodal algorithms have been explored within brain-computer interface (BCI) research, the integration of modal features often falls short of efficacy. Moreover, comprehensive multimodal classification studies employing deep learning techniques for attention state classification are limited. This paper proposes a novel EEG-fNIRS multimodal deep fusion framework (EFDFNet), which employs fNIRS features to enhance EEG feature disentanglement and uses a deep fusion strategy for effective multimodal feature integration. Additionally, we have developed EMCNet, an attention state classification network for the EEG modality, which combines Mamba and Transformer to optimize the extraction of EEG features. We evaluated our method on two attention state classification datasets and one motor imagery dataset, i.e., mental arithmetic (MA), word generation (WG) and motor imagery (MI). The results show that EMCNet achieved classification accuracies of 86.11%, 79.47% and 75.77% on the MA, WG and MI datasets using only the EEG modality. With multimodal fusion, EFDFNet improved these results to 87.31%, 80.90% and 85.61%, respectively, highlighting the benefits of multimodal fusion. Both EMCNet and EFDFNet deliver state-of-the-art performance and are expected to set new baselines for EEG-fNIRS multimodal fusion.
Individual differences and nonstationary characteristics are prominent in electroencephalography (EEG) signals. Therefore, aligning the source and target domain data becomes essential in cross-subject and cross-session classification tasks. Although many adversarial adaptation networks can achieve distribution alignment through domain-level adaptation, they tend to disregard the multimodal structure inherent in the data. In this article, we present a concise and effective adversarial paradigm for EEG emotion recognition. This approach fully utilizes the label structure information of source domain data to reuse the binary discriminator as a class-informed discriminator instead of introducing additional modules, which not only realizes domain confusion but also ensures that mode information is retained in the process of confusion to avoid mode collapse. To evaluate our method, a systematic experimental study was conducted on the public datasets SEED and SEED-IV. The average accuracy of cross-subject and cross-session scenarios achieved 90.21%, 95.47% on SEED, and 77.50%, 77.54% on SEED-IV, respectively. Compared to the existing domain adaptation methods, the evident improvements of classification performance demonstrate the feasibility and effectiveness of our method.
Background: Decoding motor intentions from electroencephalogram (EEG) signals is a critical component of motor imagery-based brain–computer interface (MI–BCIs). In traditional EEG signal classification, effectively utilizing the valuable information contained within the electroencephalogram is crucial. Objectives: To further optimize the use of information from various domains, we propose a novel framework based on multi-domain feature rotation transformation and stacking ensemble for classifying MI tasks. Methods: Initially, we extract the features of Time Domain, Frequency domain, Time-Frequency domain, and Spatial Domain from the EEG signals, and perform feature selection for each domain to identify significant features that possess strong discriminative capacity. Subsequently, local rotation transformations are applied to the significant feature set to generate a rotated feature set, enhancing the representational capacity of the features. Next, the rotated features were fused with the original significant features from each domain to obtain composite features for each domain. Finally, we employ a stacking ensemble approach, where the prediction results of base classifiers corresponding to different domain features and the set of significant features undergo linear discriminant analysis for dimensionality reduction, yielding discriminative feature integration as input for the meta-classifier for classification. Results: The proposed method achieves average classification accuracies of 92.92%, 89.13%, and 86.26% on the BCI Competition III Dataset IVa, BCI Competition IV Dataset I, and BCI Competition IV Dataset 2a, respectively. Conclusions: Experimental results show that the method proposed in this paper outperforms several existing MI classification methods, such as the Common Time-Frequency-Spatial Patterns and the Selective Extract of the Multi-View Time-Frequency Decomposed Spatial, in terms of classification accuracy and robustness.
Domain adaptation (DA) is considered to be effective solutions for unsupervised emotion recognition cross-session and cross-subject tasks based on electroencephalogram (EEG). However, the cross-domain shifts caused by individual differences and sessions differences seriously limit the generalization ability of existing models. Moreover, existing models often overlook the discrepancies among task-specific subdomains. In this study, we propose the auxiliary classifier adversarial networks (ACAN) to tackle these two key issues by aligning global domains and subdomains and maximizing subdomain discrepancies to enhance model effectiveness. Specifically, to address cross-domain discrepancies, we deploy a domain alignment module in the feature space to reduce inter-domain and inter-subdomain discrepancies. Meanwhile, to maximum subdomain discrepancies, the auxiliary adversarial classifier is introduced to generate distinguishable subdomain features by promoting adversarial learning between feature extractor and auxiliary classifier. System experiment results on three benchmark databases (SEED, SEED-IV, and DEAP) validate the model's effectiveness and superiority in cross-session and cross-subject experiments. The method proposed in this study outperforms other state-of-the-art DA, that effectively address domain shifts in multiple emotion recognition tasks, and promote the development of brain-computer interfaces.
Manifold learning with Symmetric Positive Definite (SPD) matrices has demonstrated potential for classifying Electroencephalography (EEG) in Brain-Computer Interface (BCI) applications. However, SPD matrices may lead to crucial information loss of EEG signals. This paper proposes a dimensionality reduction method based on discriminative geometric perception on the Riemannian manifold to enhance SPD matrix discriminability. Experiments on BCI Competition IV Dataset 1 and Dataset 2a show the proposed method improves accuracy by 5.0% and 19.38% respectively, demonstrating that applying discriminative geometric perception can effectively maintain robust performance associated with the dimensionality-reduced SPD matrix.
Emotion recognition based on electroencephalogram (EEG) data holds pivotal importance for advancing affective brain-computer interfaces. However, in cross-subject emotion recognition scenarios, negative transfer is likely to happen due to EEG's individual differences and inherent temporal variability. To solve these issues, this study proposes a novel domain adaptation architecture, named dual filtration subdomain adaptation network (DFSAN), to mitigate negative transfer and align subdomain features at a fine-grained category level. Firstly, the transferability of each subject was assessed to identify those with high transferability to serve as source domains. Then, with the feature alignment through subdomain metric learning, the transferable features could be obtained by dual filtration network. Finally, dual classifiers were employed to mitigate misclassifications near the decision boundary and output the recognition results. Multi-source cross-subject emotion recognition experiments were executed with SEED, SEED-IV, DEAP and SEED-V datasets, achieving recognition accuracy of 88.68 %, 67.61 %, 65.33 % and 65.57 %, respectively. Compared with other state-of-the-art domain adaptation methods, our proposed method achieved better results in cross-subject emotion recognition tasks, demonstrating the effectiveness and feasibility of DFSAN in handling negative transfer under multi-source transfer emotion recognition.
Advances in artificial intelligence have significantly enhanced intelligent assistance and rehabilitation medicine by leveraging electroencephalogram (EEG) signal recognition. Nevertheless, eliminating cross-subject variability remains a significant challenge in expending the application of EEG signal recognition to the broader society. The transfer learning strategy has been utilized to address this issue; however, multi-source domains are often treated as a single entity in transfer learning, leading to underutilization of the information from multiple sources. Furthermore, many EEG signal transfer approaches overlook the low-dimensional structural information and multivariate statistical features inherent in EEG signals, leading to inadequate interpretability and suboptimal performance. Thus, in this study, a novel multi-morphological representation approach (MMRA) was proposed for multi-source EEG signal recognition to address these issues. MMRA utilized multi-manifold mapping to extract the common invariant representation shared between the multi-source domains and target domain. It took into account the low-dimensional structure and multivariate statistical features of EEG signals to enhance the acquisition of high-quality common invariant representations. Subsequently, the multi-source domains were decomposed to extract one-to-one features. The Maximum Mean Discrepancy (MMD) loss was further applied to guide the model in obtaining high-quality private invariant representations. The performance of the proposed MMRA method was evaluated using three publicly available motor imagery datasets and a driving fatigue dataset. Experimental results demonstrated that our proposed MMRA method outperformed other state-of-the-art methods in scenarios involving multiple subjects. In conclusion, the MMRA method developed in this study can serve as a novel tool offering enhanced performance to analyze EEG signals across various subjects.
Combining the complementary advantages of electroencephalography (EEG) and functional near-infrared spectroscopy (fNIRS) shows promising potential for enhancing the decoding performance of brain-computer interfaces (BCI). However, integrating these signals through commonly used joint learning frameworks often fails to achieve the expected results. The two primary reasons are: 1) Inadequate decoupling and effective utilization of diversity features; and 2) Imbalance in the learning of multimodal features. To overcome these limitations, this paper proposes a feature decoupling and modality rebalancing framework for EEG and fNIRS integration. Specifically, a dual-decoder architecture is introduced within the standard joint learning framework. The Query decoder is employed to decouple the intermodal coupling features, while the Key decoder focuses on decoupling modality-specific features. This approach mitigates signal interference and enhances the extraction of task-relevant information. Furthermore, we introduce a Gradient Rebalancing (GradReb) Strategy to regulate gradient flow during training, ensuring balanced learning across both modalities and mitigating the risk of one modality predominating the process. Experimental results demonstrate that our method effectively leverages the complementary information from both EEG and fNIRS, significantly improving the performance of the joint learning framework. The proposed joint framework provides a scalable solution for hybrid BCI systems and offers new insights for other multimodal applications.
Graph Convolutional Networks (GCNs) have shown promise in motor imagery electroencephalogram (EEG) signals classification by modeling spatial dynamics and brain connectivity. However, over-smoothing remains a challenge, leading to homogenized node features and reduced discrimination. To address this, we propose an Adaptive Sparse Awareness-Spatiotemporal Graph Convolutional Network (ASA-STGCN) that combines adaptive sparse graph convolution with attention mechanisms. Notably, a Graph Sparse Convolutional Network (GSCN) in the Adaptive Sparse Awareness Spatial Module (ASAM) enhances brain region feature selection, while the Graph Node Neighborhood Awareness Layer (GNNAL) applies self-attention to reinforce critical topological relationships. The Multi-scale Temporal Convolution Module (MTCM) captures both transient and sustained temporal dependencies. Experimental results achieve accuracies of 97.2%±3.4% (binary) and 83.6%±4.9% (four-class) on BCIC-IV-2a, 96.6%±3.1% (binary) on BCIC-III IVa, and 83.41%±4.3 (binary) on OpenBMI. Discussion confirms the model's effectiveness and its potential to support EEG-based neurorehabilitation and clinical brain computer interface applications.
Owing to individual difference, it is challenging to decode the target subject's mental intentions by applying existing models in cross-subject Brain-Computer Interface tasks. The transfer learning methods have shown promising performance in this field, but they still suffer from poor presentations on temporal correlations characteristics of EEG signals. This paper proposes a multi-view convolution-transformer based domain adaptation framework for cross-subject motor imagery classification tasks. Firstly, to exploit the frequency diversity of EEG signals, we decomposed EEG signals into several overlapping frequency views and extracted frequencyrelated spatial and temporal features by parallel spatiotemporal convolution block. Subsequently, we use transformer blocks to extract long-range dependencies and narrow the marginal distribution between source and target domains. Eventually, a classifier and a domain discriminator were used for domain adaptation, and a mixed loss was employed to align conditional distributions. We conducted model validation on the BCI Competition IV 2a and 2b datasets and achieved average accuracies of 77.8 % and 80.1 %, respectively. The experimental results show that our proposed framework outperforms traditional deep adversarial domain adaptive methods.