Drug-resistant epilepsy (DRE) affects millions of people worldwide and remains a major therapeutic challenge, largely due to the difficulty in precisely localizing the epileptogenic zone (EZ). Current electrophysiological biomarkers often lack robustness across ictal-interictal states and clinical centers. Furthermore, neuronal excitation-inhibition (E/I) dynamics-a key pathophysiological mechanism-has not yet been systematically characterized for clinical translation. In this multicenter retrospective study, we introduce a computationally grounded framework leveraging stereoelectroencephalography (SEEG) -derived 1/f spectral signatures as a proxy for E/I ratio to enable machine learning-guided EZ identification. We analyzed SEEG recordings capturing a total of 122 seizures from 38 patients with DRE who achieved complete seizure freedom (Engel Class I) post-surgery. Cohort-level analyses revealed that the EZ exhibited a more negative E/I ratio compared to non-epileptogenic zones (NEZ) across both interictal and ictal states ( ${p} \lt 0.001$ , after FDR correction), indicating a significant imbalance of E/I dynamics. These findings were corroborated at the individual level, where 84.2% (32/38) and 68.4% (26/38) of patients showed significant EZ-NEZ separation during ictal and interictal periods, respectively ( ${p} \lt 0.05$ , after FDR correction). This discriminative capacity was consistent across surgical modalities (resection/ablation) and clinical centers. By incorporating E/I dynamics and multi-band average power spectral density (PSD) as features to train 11 machine learning models (e.g. SVM, Random Forest), we found that the Random Forest classifier achieved 0.84 accuracy (AUC = 0.90) in EZ localization, demonstrating robust generalizability. The study established the E/I dynamics as a clinically translatable generalizable framework for refining surgical targeting in DRE. To promote reproducibility and community validation, the implementation code is publicly available at: https://github.com/wyl1994/Source-code-and-dataset/tree/main.
Existing deep learning-based automatic sleep staging methods for neonates predominantly rely on CNN or RNN architectures, often overlooking information on inter-channel spatial and inter-stage temporal dynamics. Although GNN-based and graph-based methods can leverage such information, they solely focus on functional connectivity from linear relationships instead of incorporating effective connectivity from information flow. Thus, a multi-view self-attention based on the brain network method is proposed for neonatal sleep staging (MABNSleepNet), which consists of a Brain Network Extractor (BNE), a Context-attaching Module (CAM), a Multi-view Self-attention Module (MSAM), and a Feature-aggregating Module (FAM). The proposed method leverages multi-view self-attention and brain connectivity analysis to comprehensively capture both spatial, temporal, and frequency-domain features, while incorporating inter-stage temporal context for accurate neonatal sleep staging. It was validated on a clinical neonatal sleep dataset from the Children’s Hospital of Fudan University (CHFU) with 64 eight-channel EEG recordings. Employing the subject-wise 10-fold validation method with six testing set and 58 training set, the proposed method achieved impressive results on two-stage (wakefulness and sleep) and three-stage (wakefulness, active sleep, and quiet sleep) tasks respectively: with accuracies of 0.920 and 0.862, F1-scores of 0.909 and 0.861, as well as Cohen’s Kappa coefficients of 0.817 and 0.793. Experiment results demonstrate the proposed method’s superiority over most current state-of-the-art methods in neonatal sleep staging, providing a novel avenue for in-depth analysis of sleep staging from the brain connectivity perspective.
Dexterous hand motor functions are highly flexible and finely controlled by complex neural commands from the motor cortex. However, in patients with brain injuries such as stroke, restoring fine motor control from the perilesional cortex remains extremely challenging. A major obstacle is the absence of appropriate non-human primate models to elucidate the behavioral and neural signatures of hand motor function during recovery following treatments. Here, we present a new non-human primate model that reflects the motor function recovery processes following lesion-induced hand paralysis after contralateral C7 nerve transfer (CC7) surgery, which establishes a new neural pathway from the ipsilateral cortex to control the paralyzed hand. By developing a hand reach-to-pinch task and quantifying finger kinematics, we established systematic, objective profiles of fine motor recovery in human patients and monkey models following CC7 treatment. Furthermore, when considering behavioral aspects, spontaneous recovery of hand motor skills was notably limited in human patients and monkey models, as indicated by the consistently abnormal “thumb-in-palm” patterns observed in finger kinematic analysis. However, the CC7 surgery gradually restored the finger kinematic patterns during hand-pinch actions to nearly identical patterns to those of the healthy hand. In addition, the human functional MRI and macaque electrophysiology results revealed, on a neural level, the emergence of a new command area and its spiking-based motor-command refinements specifically for the paralyzed hand in the contralesional M1 and premotor cortex (PMC) after CC7 treatment. Thus, our findings strongly support the notion that modifying peripheral nerve pathways greatly promotes the recovery of dexterous motor function in a paralyzed hand by reconstructing new motor-control neural mechanisms within the ipsilateral healthy motor cortex.
Background: Agent skills are increasingly deployed as modular, reusable capability units in AI agent systems. Medical research agent skills require safeguards beyond general-purpose evaluation, including scientific integrity, methodological validity, reproducibility, and boundary safety. This study developed and preliminarily evaluated a domain-specific audit framework for medical research agent skills, with a focus on reliability against expert review. Methods: We developed MedSkillAudit (skill-auditor@1.0), a layered framework assessing skill release readiness before deployment. We evaluated 75 skills across five medical research categories (15 per category). Two experts independently assigned a quality score (0-100), an ordinal release disposition (Production Ready / Limited Release / Beta Only / Reject), and a high-risk failure flag. System-expert agreement was quantified using ICC(2,1) and linearly weighted Cohen's kappa, benchmarked against the human inter-rater baseline. Results: The mean consensus quality score was 72.4 (SD = 13.0); 57.3
Accurate segmentation of lumbar paraspinal muscles from clinical T2-weighted MRI is essential for quantitative assessment of muscle atrophy and fatty infiltration, yet remains challenging due to the anisotropic nature of clinical scans with high in-plane resolution but sparse axial sampling, severe fatty infiltration causing indistinct boundaries, and the scarcity of finely annotated datasets. To address these challenges, we present a large-scale benchmark dataset for Lumbar Paraspinal Muscle Segmentation (LPMSeg), comprising 260 patients with 2,340 expert-annotated slices spanning L3–S1 levels, providing high-resolution separate masks for multifidus and erector spinae muscles. Furthermore, we propose S3Net, a hybrid convolutional-transformer architecture built upon a Compress-Model-Reconstruct paradigm that employs strategic axial downsampling to reduce quadratic complexity while preserving anatomical continuity. A novel Spine Consistency Loss is introduced to penalize abrupt probability variations, effectively bridging fat-induced gaps. Extensive experiments demonstrate that S3Net achieves state-of-the-art performance, while requiring fewer parameters.
Drug-resistant epilepsy (DRE) affects approximately 30% of epilepsy patients, with surgical cure rates below 70%. This challenge drives a fundamental paradigm shift from localizing a discrete epileptogenic zone toward characterizing and modulating the dysfunctional brain networks that initiate and propagate seizures. This review critically synthesizes how computational neuroelectrophysiology and artificial intelligence (AI) converge to propel this shift. We chart the trajectory from quantifying local biomarkers, including high-frequency oscillations and excitation-inhibition balance metrics, to mapping the topological properties of epileptic networks through functional and effective connectivity. The integration of these network features with AI techniques, particularly spatiotemporal deep learning architectures, has demonstrated significant potential for enhancing both the localization of the epileptogenic network and the prediction of postoperative outcomes. However, the translation of this network-centric paradigm into clinical practice remains constrained by several challenges, including spatial sampling bias, the inherent label instability of surgical outcomes, and the 'black-box' interpretability crisis. Future directions must emphasize EEG foundation models, causal AI, and the development of interactive multimodal AI agents. The convergence of mechanism-driven network models and data-driven AI frameworks promises a new era of personalized, network-guided therapy for DRE.
Employing a minimal array of electroencephalography (EEG) channels for neonatal sleep stage classification is essential for data acquisition in the Internet of Medical Things (IoMT), as single-channel and edge-based features can reduce data transfer and processing requirements, enhancing cost-effectiveness and practicality. In this paper, we evaluate the efficacy of a single channel and the viability of a binary classification scheme for discerning awake and sleep states and transitions to quiet sleep. For this, two datasets of EEG signals for neonate sleep analysis were recorded from Children's Hospital of Fudan University, Shanghai, comprising recordings from 64 and 19 neonates, respectively. From each epoch, a diverse ensemble of 490 features was extracted through a blend of discrete and continuous wavelet transforms (DWT, CWT), spectral statistics, and temporal features. In addition, we introduced an innovative hybrid univariate and ensemble feature selection approach with multidomain feature fusion, a stacking-based ensemble classifier that outperforms existing work. We achieved 90.37%, 91.13%, and 94.88% accuracy for sleep/awake, quiet sleep/non-quiet sleep, and quiet sleep/awake, respectively. This was corroborated by significant Kappa values of 77.5%, 80.29%, and 89.76%. Using SelectPercentile, we devised three distinct feature selection mechanisms: one using DWT, one with CWT, and another incorporating both spectral and temporal features. Subsequently, SelectKBest was used to determine the most effective features. For our stacked model, we incorporated a trifecta of the ExtraTree model with variable estimators, a Random Forest, and an Artificial Neural Network (ANN) as base classifiers, and for the final prediction phase, ANN was implemented again. The model's performance was evaluated using K-fold and leave-one-subject cross-validation.
Electrocardiography(ECG)-based detection methods offer a promising alternative to traditional polysomnography (PSG) for the diagnosis of sleep apnea. However, existing ECG-based methods remain limited by insufficient extraction of physiologically meaningful features and inadequate modeling of heterogeneous ECG-derived representations, which can restrict their detection performance. We propose PASE-MST, a multi-stream framework that jointly uses raw ECG, R-wave amplitude (RA), RR interval (RRI), RR-interval derivative (RRID), and an enhanced cardiopulmonary coupling (CPC) sequence. Dedicated streams with pre-activation (PA) residual blocks process different input types, and squeeze-and-excitation (SE)-based fusion weights complementary features across streams. Within the framework, the CPC computation is further adapted by replacing QRS-amplitude-based ECG-derived respiration with an R-wave-slope-based surrogate. We evaluated PASE-MST on the public Apnea-ECG database and an independent clinical dataset from Huashan Hospital. The results show that PASE-MST achieves better performance than recent ECG-based obstructive sleep apnea (OSA) detection methods. On Apnea-ECG, the proposed method achieved an accuracy of 93.13±0.39%, a sensitivity of 91.78±1.57%, and a specificity of 93.97±0.77%. On the Huashan Hospital Database (HHD), it achieved an accuracy of 93.32±0.25%, a sensitivity of 89.29±1.45%, and a specificity of 94.18±0.65%. These results demonstrate that the proposed framework achieves consistent OSA detection performance under subject-independent evaluation in both a public benchmark and an independent clinical cohort.
The development of wearable sensor-based sign language recognition systems has become a solution to facilitate effective communication among hearing-impaired groups, but achieving high integration, sign language standardization, and anti-environmental interference remains challenging. Here, we design a smart glove system for real-time sign language interpretation based on a composite foam with a cross-dimensional conductive network, integrating flexible switches, pressure sensors, custom miniaturized circuits, and deep learning modules. The sensor exhibits electromagnetic shielding, thermal management, and antibacterial capabilities, enhancing the smart glove's adaptability to the external environment. The elastic conductive framework of the foam allows the system to realize the start/stop function and fast response to gestures. In addition, a deep learning model of multi-component collaboration and mechanism fusion is constructed, along with the establishment of a comprehensive set of sign language rules, which realizes 99.4 % accurate recognition of 26 letters through only three pressure sensors. Overall, our proposed strategy provides a new way for smart gloves to work stably in harsh environments, and is expected to eliminate communication barriers among hearing-impaired groups due to sign language diversity and cultural differences.
Epilepsy is a neurological disorder characterized by abnormal discharges in the brain, which can occur at any age. Seizures not only have the potential to cause cognitive impairments but also lead to anxiety, depression, and social isolation, severely impacting the quality of life for patients. Therefore, research on epilepsy detection across all-age groups is of significant importance. Electroencephalography (EEG) is a commonly used tool for diagnosing epilepsy, as analyzing EEG signals can help determine seizure events. Existing epilepsy detection models typically rely on EEG data from a single age group, limiting their generalization capabilities and applicability. This study proposes a text-guided model for epilepsy detection in patients across all ages, involving four publicly available datasets that encompass neonates, children, and adults.The APTGNet model uses age prompts combined with textual contrastive learning to train the model, achieving epilepsy classification across all-age groups. The model consists of a text encoder and an EEG encoder, where the text encoder converts text into high-dimensional semantic features, and the EEG encoder utilizes a Dynamic Attention Module (DAM) to dynamically focus on EEG signals from patients of different age groups. By calculating the cosine similarity between text and EEG features, APTGNet can accurately classify seizure signals, demonstrating strong generalization capabilities and clinical application potential. Our model exhibits good performance across multiple datasets for all-age groups, particularly in the following aspects: (1) strong overall performance in seizure detection datasets across various age groups; (2) effective performance and generalization enabled by text guidance and age prompts; and (3) unifying seizure detection tasks across multiple age groups through the DAM.
Obstructive sleep apnea (OSA) and sleep fragmentation are closely linked physiological phenomena that play crucial roles in the diagnosis and management of sleep disorders. While numerous deep learning models have been developed for either OSA detection or sleep stage classification, few attempts have been made to address both tasks simultaneously. To this end, we propose MT-TASPPNet (Multi-Task Triple Atrous Spatial Pyramid Pooling Network), a unified multi-modal multi-task network that jointly performs automatic OSA event detection and sleep staging. The model integrates modality-specific feature extractors for EEG, ECG, and airflow signals, and employs Atrous Spatial Pyramid Pooling modules in both the modality-specific and shared representation pathways to capture multi-scale temporal-frequency patterns. Additionally, an EOG-guided prior mechanism is incorporated to enhance the discrimination of subtle sleep stages. We use a 3-min input window (1-min target with $\pm$ 1-min context) and evaluate our method on three large-scale datasets: SHHS1, SHHS2, and Sydney Sleep Biobank. The model achieves OSA detection accuracy between 0.798 and 0.884 (MF1: 0.772 to 0.821), and sleep staging accuracy between 0.776 and 0.834 (MF1: 0.735 to 0.749, $\mathcal {K}$: 0.697 to 0.77). Notably, the model maintains consistent performance despite data heterogeneity and individual variability. These results validate the stability and adaptability of MT-TASPPNet in clinical settings, paving the way for efficient and scalable multi-task sleep analysis systems.
IntroductionBradykinesia, a cardinal physical dysfunction of Parkinson's disease (PD), is generally evaluated by Section III of the Movement Disorders Society-sponsored revision of the unified Parkinson's disease rating scale (MDS-UPDRS). The evaluation process requires the supervision of clinicians; therefore, the results may be subjective and increase clinicians' workload.MethodsTo compensate for these drawbacks, this study proposes a task-specific machine learning system that incorporates wearable sensors to automatically monitor bradykinesia. Initially, this study enhanced the peak detection algorithm by considering specific motion characteristics, such as movement amplitude, to make it more adaptable to individualized signals. Furthermore, a task-specific score prediction system was proposed, incorporating the optimal sensor placement positions and appropriate score prediction algorithms. Specifically, seven machine learning models and four ensemble methods were explored for each of the five MDS-UPDRS III tasks.ResultsThe system was tested in 21 patients and eight control individuals. The performance of the system was evaluated under three scenarios: precise prediction (0 vs. 1 vs. 2 vs. 3), abnormal/normal [0 vs. (1, 2, 3)], and normal-moderate/severe [(0,1,2) vs. 3], with the highest average F1 score reaching 0.8722 and the lowest total root mean square error at 0.3214 among tasks in critical status prediction. Furthermore, this study, through statistical analysis, suggested specific scoring features tailored to each task.DiscussionThis study demonstrated the feasibility of accurately and automatically monitoring bradykinesia of PD.
Background: Sleep-disordered breathing (SDB) is frequently accompanied by autonomic nervous system (ANS) dysfunction, which is closely associated with an increased incidence of cardiovascular diseases and elevated mortality risk. Heart rate variability (HRV) serves as a classic metric for evaluating sympathovagal balance; however, the specific impacts of four distinct types of respiratory events-obstructive apnea (OA), central apnea (CA), mixed apnea (MA), and hypopnea (HYP)-on HRV remain underinvestigated. Utilizing ultra-short-term HRV analysis, this study aimed to evaluate the immediate effects of different respiratory events on ANS function, while further exploring the modulatory roles of arousal, Apnea-Hypopnea Index (AHI) severity and sleep stages (non-rapid eye movement [NREM] vs. rapid eye movement [REM]). Methods: A total of 108 patients with SDB undergoing overnight polysomnography (PSG) were included. A total of 19,862 respiratory events, including obstructive apnea (OA), central apnea (CA), mixed apnea (MA), and hypopnea (HYP), were analyzed using 15 s ECG segments. Linear mixed-effects models (LMMs) and estimated marginal means (EMMs) with Sidak-adjusted pairwise comparisons were constructed to evaluate differences in ECG-derived features and to analyze differences between event types. Results: Central apnea (CA) was associated with significantly reduced HRV and heart rate indices, including Standard Deviation of Successive Differences (SDSD), Root Mean Square of the Successive (RMSSD), Standard Deviation 1 (SD1), and heart rate (HR), compared with other respiratory event types (all p < 0.05). Across all event types, HRV metrics exhibited consistent dynamic changes before, during, and after respiratory events (all p < 0.001), characterized by a decrease during the event followed by post-event recovery. In the interaction effect of sleep stage, SDSD was significantly lower in CA compared with both OA (estimate = -11.67, 95% CI -18.78 to -4.59, p < 0.001) and HYP (estimate = -11.38, 95% CI -18.55 to -4.20, p < 0.001) during NREM sleep. No significant differences in HRV parameters, heart rate, or QRS duration were observed between OA and HYP (all p > 0.05). Conclusions: This study is the first to elucidate the differential impacts of four distinct types of sleep respiratory events on ultra-short-term HRV, confirming that CA events exert the most profound effects on autonomic function. These findings suggest that the proportion of CA occurrences could serve as a more precise biomarker for identifying individuals at high risk for cardiovascular diseases within the SDB population.
IntroductionSleep apnea and hypopnea syndrome (SAHS) is a prevalent disorder with profound adverse effects on health and overall quality of life, thereby necessitating the development of accurate and accessible screening tools. Electrocardiogram (ECG)-based analysis, being non-invasive and readily deployable in low-cost hardware, offers a particularly convenient approach for SAHS screening and preliminary diagnosis. However, conventional time-frequency analysis often fails to capture the subtle yet critical patterns in ECG signals due to the Heisenberg uncertainty principle, leading to limited resolution and information loss.MethodsTo overcome these limitations, this study proposes a Dual Stream Cross Attention Fusion Network (DSCAFNet) based on the uncertainty-mitigated time-frequency representations generated via Synchrosqueezing Transform (SST). The framework uniquely constructs two complementary, high-fidelity SST-based representations, which are strategically designed to provide distinct yet synergistic perspectives on the complex, non-stationary dynamics of SAHS. A dedicated cross-attention fusion module then harnesses these complementary views, enabling the model to discriminatively integrate multi-resolution features for significantly enhanced pattern recognition.ResultsExtensively evaluated on the public Apnea-ECG dataset, DSCAFNet achieves an accuracy of 0.9572, a sensitivity of 0.9575, a specificity of 0.9584, and an F1-score of 0.9557, performing on par with state-of-the-art methods. More importantly, rigorous validation on a private Huashan-apnea dataset yields an accuracy of 0.9003 for binary classification and 0.7564 for four-class subtyping, demonstrating strong effectiveness and generalization.ConclusionThese consistent results across datasets highlight DSCAFNet as a promising framework for intelligent and accessible SAHS screening, with potential for integration into portable data acquisition systems combined with cloud-based analysis.
Stereoelectroencephalography (SEEG)-guided radiofrequency thermocoagulation is the mainstream treatment for drug-resistant epilepsy (DRE), yet non-invasive patient-specific localization of potential epileptogenic zone (EZ) prior to SEEG electrode implantation remains a critical unmet clinical need, hindered by limited automation, suboptimal accuracy, and poor cross-patient generalizability. To address these gaps, we developed an automated non-invasive EZ localization framework that integrates scalp EEG source imaging, high-resolution time-frequency analysis, and deep learning. This multicenter retrospective study included 97 seizure episodes from 37 surgically confirmed DRE patients with Engel Class I post-operative seizure freedom. Three time-frequency spectrograms were generated based on the source-reconstructed signal, and fed into four deep learning architectures (ResNet, VGG, DenseNet, Swin Transformer) for channel-level EZ binary classification with patient-level leave one-out cross-validation to validate personalized localization for unseen patients. The ResNet18-Superlet combination achieved optimal performance (accuracy = 81.15% ± 4.83%, AUC = 84.82% ± 5.27%) in the adult cohort, with robust generalization to pediatric and mixed cohorts, significantly outperforming conventional high-frequency oscillation (HFO)-based methods. This framework enables accurate personalized pre-surgical EZ map ping to optimize SEEG implantation planning, with open-source code available at https://github.com/wyl1994/Source-code-for Patient-Specific-Non-Invasive-Epileptogenic-Zone-Localization to ensure reproducibility and clinical translation.
Background and objective: Sleep analysis provides an important window into neonatal neurodevelopment, but automatic neonatal sleep staging remains challenging because neonatal EEG is nonstationary, low in signal-to-noise ratio, and strongly affected by inter-channel spatial interactions. Existing approaches often rely on either handcrafted connectivity features or purely data-driven deep learning features, and therefore may not fully exploit complementary functional-connectivity and adaptive spatial information. This study proposes HSleepNet, a hybrid adaptive brain network architecture for neonatal sleep staging. Methods: HSleepNet integrates three key components: (i) a multi-dimensional brain network (MDBN) branch that extracts Pearson correlation, mutual information, phase-locking value, phase-lag index, and phase-slope-index connectivity; (ii) an adaptive brain network (ABN) branch that learns row-normalized channel relationships from CNN–BiLSTM features; and (iii) a spatial feature learning module that fuses connectivity priors and adaptive graph representations. The model was evaluated using subject-wise 10-fold cross-validation on the private CHFD neonatal EEG dataset comprising 64 recordings. MASS-S3 was additionally used as an adult cross-population benchmark rather than as evidence of external neonatal validation. Results: On CHFD, HSleepNet achieved an accuracy of 81.4% and an MF1 score of 0.813 for the three-stage wakefulness–AS–QS task. On MASS-S3, the model achieved an accuracy of 85.6% and an MF1 score of 0.791 for the five-stage adult sleep task. These results suggest that the hybrid MDBN–ABN design provides complementary information for multi-channel EEG sleep staging. Conclusions: HSleepNet provides preliminary evidence that combining domain-informed connectivity features with adaptive graph learning can improve neonatal sleep staging on the studied cohort. The main limitation is that neonatal validation was performed on a single private clinical dataset; therefore, broader external neonatal validation and more extensive clinical outcome analysis are required before making claims about deployment or general clinical utility.
Sleep spindles and K-complexes in electroencephalo gram (EEG) signals are brief yet physiologically essential signatures of the sleeping brain, supporting memory consolidation, arousal regulation, and neurological assessment. However, existing detection methods often face severe class imbalance, temporal inconsistencies, and uncertain event boundaries. We proposed a unified, physiologically informed framework that integrates data preprocessing, multimodal feature extraction, temporal modeling, and event-level refinement. Context-aware negative filtering and segment-level oversampling effectively alleviated severe class imbalance; multimodal feature combination captured both spectral and morphological cues; a Bidirectional Long Short-Term Memory (BiLSTM) backbone enhanced by segment level dropout and bidirectional hard example mining improves robustness under imbalance; and a physiologically plausible post-processing enforced temporal coherence on the predicted sequences. Evaluations on the publicly available polysomnog raphy datasets DREAMS and MASS demonstrated consistent 3-10% F1-score improvements over state-of-the-art baselines. The proposed framework attained event-level F1-scores of 0.796 (sleep spindles on DREAMS), 0.938 (K-complexes on DREAMS), 0.882 (sleep spindles on MASS), and 0.903 (K-complexes on MASS), substantially outperforming recent deep-learning-based detectors.
As a critical factor in diagnostic work-up and treatment decision-making process of sleep-related breathing disorders, accurate localization of obstructive sites in the upper airway is in dire need. Snoring, as a dynamic acoustic signal, carries informative information relating to the sites and degree of obstruction in the upper airway, offering a non-invasive, cost-effective solution for obstructive sites recognition. However, most of existing snoring-based methods for recognizing obstructive sites only involve limited information (either mainly concentrated on traditional acoustic characteristics or spectrogram features), which may omit dynamic pathological information. Moreover, existing methods proceed from either a one-dimensional (1D) signal or two-dimensional (2D) image perspective, where complementary information from the other modality may be overlooked. In this paper, a multi-modal framework, which combines 1D snoring waveform and 2D Composite Acoustic Feature Graph (CAF-Graph), is proposed. 1D snoring waveform perceives fine time structure and local patterns, aiming at learning high-level discriminative representations by neural networks. 2D CAF-Graph is dedicated to emphasizing dynamic spatio-temporal and physiological-acoustic characteristic of snoring, which concatenates acoustic features related to Prosodic, Formant, Spectral, and Cepstral characteristics. Further, a multi-modal fusion network (BMFNet) effectively integrates independent and interactive information between single-modal features, which offers a more comprehensive perspective. The recognition task was formulated as a three-class classification problem, including upper (snoring caused by upper-level obstruction), lower (snoring caused by lower-level obstruction), and silence (obstruction without snoring). The proposed method was validated on a clinical dataset collected in the ENT institute and Department of Otorhinolaryngology, Eye & ENT Hospital, Fudan University, where reached 81.2% Accuracy, 86.8% Weighted Average Precision, 81.2% Weighted Average Recall, and 82.3% Weighted Average F1-Score. Results exhibit the effectiveness of multi-modal feature representations for snoring, providing a novel insight for obstructive sites recognition tasks.
Sleep staging classification holds significant clinical importance in the diagnosis of sleep-related disorders. Sequence-to-sequence and sequence-to-epoch models capture long-range temporal dependencies and typically outperform epoch-by-epoch methods, but their reliance on multi-epoch EEG inputs leads to higher computational costs and processing latency, hindering their application in real-time wearable systems. To address this issue, we propose Sequence-to-Epoch Knowledge Distillation (S2EKD), which transfers temporal knowledge from a high-latency sequence teacher to a single-epoch student, largely preserving accuracy while enabling low-latency, device-friendly inference. The core of our framework is a novel heterogeneous-input knowledge distillation process. A high-performance, sequence-based teacher network processes a full sequence of Tepoch inputs, while a lightweight student network is trained using only the corresponding middle epoch, Sepoch (or Tepoch, middle). By distilling sequence-level context from the teacher, the student makes temporally informed predictions while preserving its single-epoch feature-extraction pathway. This knowledge transfer is enabled by our novel Seq2Epoch Weight (S2EW) module. Guided by S2EW position-aware weights, the student model learns to align its intermediate features with those of the context-rich teacher, ultimately aggregating its own subepoch features to produce a single, contextually aware prediction. We evaluated our method on three public sleep datasets (Sleep-EDF20, MASS-SS3, and ISRUC-S3). The results demonstrate that S2EKD significantly enhances the performance of single-epoch models over existing methods by successfully transferring sequential knowledge. Crucially, our framework reduces model size to just 42.6% of the teacher network while simultaneously improving performance by 2-3% and eliminating the input latency inherent to sequence-based approaches, clearing a key barrier for real-time deployment. Furthermore, its versatility was confirmed by successfully adapting it to diverse sequence-based teacher models.