Rapid-eye-movement (REM) sleep behaviour disorder (RBD) is a primary sleep disorder strongly associated with Parkinson's disease. Assessing sleep structure in RBD is important for understanding the underlying pathophysiology and developing diagnostic methods. However, the performance of automated sleep stage classification (ASSC) models is considered suboptimal in RBD, for both models utilising neurological signals ("ExG": EEG, EOG, and chin EMG) and heart rate variability combined with body movements (HRVm). Here, we explore this underperformance through the categorical representation of sleep macrostructure (i.e., hypnogram) and a representation that leverages the underlying probability distribution of ASSCs (i.e., hypnodensity). By comparing the RBD population (n = 36) to a sex- and age-matched group of OSA patients chosen for their anticipated similarly decreased sleep stability, we confirm lower 4-stage classification performance in both ExG-based ASSC (RBD: κ = 0.74, OSA: κ = 0.80) and HRVm-based ASSC (RBD: κ = 0.50, OSA: κ = 0.63). Stages showing lower agreement in RBD, namely, N1 + N2 and REM sleep, exhibited elevated ambiguity in the hypnodensity, indicating more ambiguous classification distributions. Limited differences in bout durations between RBD and OSA suggested sleep instability is not necessarily driving lower agreement in RBD. However, stage transitions in OSA showed more abrupt changes in the underlying probability distribution, while RBD transitions had a more continuous profile, possibly complicating classification. Although both ExG-based and HRVm-based automated sleep staging in RBD remain challenging, hypnodensity analysis is informative for the characterisation of (RBD) sleep and can capture potential drivers of classification disagreement.
Objective: wearable sensor technology has progressed significantly in the last decade, but its clinical usability for the assessment of obstructive sleep apnea (OSA) is limited by the lack of large and representative datasets simultaneously acquired with polysomnography (PSG). The objective of this study was to explore the use of cardiorespiratory signals common in standard PSGs which can be easily measured with wearable sensors, to estimate the severity of OSA. Methods: an artificial neural network was developed for detecting sleep disordered breathing events using electrocardiogram (ECG) and respiratory effort. The network was combined with a previously developed cardiorespiratory sleep staging algorithm and evaluated in terms of sleep staging classification performance, apnea-hypopnea index (AHI) estimation, and OSA severity estimation against PSG on a cohort of 653 participants with a wide range of OSA severity. Results: four-class sleep staging achieved a kappa of 0.69 versus PSG, distinguishing wake, combined N1-N2, N3 and REM. AHI estimation achieved an intraclass correlation coefficient of 0.91, and high diagnostic performance for different OSA severity thresholds. Conclusions: this study highlights the potential of using cardiorespiratory signals to estimate OSA severity, even without the need for airflow or oxygen saturation (SpO2), traditionally used for assessing OSA. Significance: while further research is required to translate these findings to practical and unobtrusive sensors, this study demonstrates how existing, large datasets can serve as a foundation for wearable systems for OSA monitoring. Ultimately, this approach could enable long-term assessment of sleep disordered breathing, facilitating new avenues for clinical research in this field.
Overnight sleep staging is an important part of the diagnosis of various sleep disorders. Polysomnography is the gold standard for sleep staging, but less-obtrusive sensing modalities are of emerging interest. Here, we developed and validated an algorithm to perform “proxy” sleep staging using cardiac and respiratory signals derived from a chest-worn accelerometer. We collected data in two sleep centers, using a chest-worn accelerometer in combination with full PSG. A total of 323 participants were analyzed, aged 13–83 years, with BMI 18–47 kg/m2. We derived cardiac and respiratory features from the accelerometer and then applied a previously developed method for automatic cardio-respiratory sleep staging. We compared the estimated sleep stages against those derived from PSG and determined performance. Epoch-by-epoch agreement with four-class scoring (Wake, REM, N1+N2, N3) reached a Cohen’s kappa coefficient of agreement of 0.68 and an accuracy of 80.8%. For Wake vs. Sleep classification, an accuracy of 93.3% was obtained, with a sensitivity of 78.7% and a specificity of 96.6%. We showed that cardiorespiratory signals obtained from a chest-worn accelerometer can be used to estimate sleep stages among a population that is diverse in age, BMI, and prevalence of sleep disorders. This opens up the path towards various clinical applications in sleep medicine.
Automatic sleep staging based on cardiorespiratory signals from home sleep monitoring devices holds great clinical potential. Using state-of-the-art machine learning, promising performance has been reached in patients with sleep disorders. However, it is unknown whether performance would hold in individuals with potentially altered autonomic physiology, for example under the influence of medication. Here, we assess an existing sleep staging algorithm in patients with sleep disorders with and without the use of beta blockers. We analyzed a retrospective dataset of sleep recordings of 57 patients with sleep disorders using beta blockers and 57 age-matched patients with sleep disorders not using beta blockers. Sleep stages were automatically scored based on electrocardiography and respiratory effort from a thoracic belt, using a previously developed machine-learning algorithm (CReSS algorithm). For both patient groups, sleep stages classified by the model were compared to gold standard manual polysomnography scoring using epoch-by-epoch agreement. Additionally, for both groups, overall sleep parameters were calculated and compared between the two scoring methods. Substantial agreement was achieved for four-class sleep staging in both patient groups (beta blockers: kappa = 0.635, accuracy = 78.1
Summary Automatic estimation of sleep structure is an important aspect in moving sleep monitoring from clinical laboratories to people's homes. However, the transition to more portable systems should not happen at the expense of important physiological signals, such as respiration. Here, we propose the use of cardiorespiratory signals obtained by a suprasternal pressure (SSP) sensor to estimate sleep stages. The sensor is already used for diagnosis of sleep‐disordered breathing (SDB) conditions, but besides respiratory effort it can detect cardiac vibrations transmitted through the trachea. We collected the SSP sensor signal in 100 adults (57 male) undergoing clinical polysomnography for suspected sleep disorders, including sleep apnea syndrome, insomnia, and movement disorders. Here, we separate respiratory effort and cardiac activity related signals, then input these into a neural network trained to estimate sleep stages. Using the original mixed signal the results show a moderate agreement with manual scoring, with a Cohen's kappa of 0.53 in Wake/N1–N2/N3/rapid eye movement sleep discrimination and 0.62 in Wake/Sleep. We demonstrate that decoupling the two signals and using the cardiac signal to estimate the instantaneous heart rate improves the process considerably, reaching an agreement of 0.63 and 0.71. Our proposed method achieves high accuracy, specificity, and sensitivity across different sleep staging tasks. We also compare the total sleep time calculated with our method against manual scoring, with an average error of −1.83 min but a relatively large confidence interval of ±55 min. Compact systems that employ the SSP sensor information‐rich signal may enable new ways of clinical assessments, such as night‐to‐night variability in obstructive sleep apnea and other sleep disorders.
Introduction: The apnea-hypopnea index (AHI), defined as the number of apneas and hypopneas per hour of sleep, is still used as an important index to assess sleep disordered breathing (SDB) severity, where hypopneas are confirmed by the presence of an oxygen desaturation or an arousal. Ambulatory polygraphy without neurological signals, often referred to as home sleep apnea testing (HSAT), can potentially underestimate the severity of sleep disordered breathing (SDB) as sleep and arousals are not assessed. We aim to improve the diagnostic accuracy of HSATs by extracting surrogate sleep and arousal information derived from autonomic nervous system activity with artificial intelligence. Methods: We used polysomnographic (PSG) recordings from 245 subjects (148 with simultaneously recorded HSATs) to develop and validate a new algorithm to detect autonomic arousals using artificial intelligence. A clinically validated auto-scoring algorithm (Somnolyzer) scored respiratory events, cortical arousals, and sleep stages in PSGs, and provided respiratory events and sleep stages from cardio-respiratory signals in HSATs. In a four-fold cross validation of the newly developed algorithm, we evaluated the accuracy of the estimated arousal index and HSAT-derived surrogates for the AHI. Results: The agreement between the autonomic and cortical arousal index was moderate to good with an intraclass correlation coefficient of 0.73. When using thresholds of 5, 15, and 30 to categorize SDB into none, mild, moderate, and severe, the addition of sleep and arousal information significantly improved the classification accuracy from 70.2% (Cohen's κ = 0.58) to 80.4% (κ = 0.72), with a significant reduction of patients where the severity category was underestimated from 18.8% to 7.3%. Discussion: Extracting sleep and arousal information from autonomic nervous system activity can improve the diagnostic accuracy of HSATs by significantly reducing the probability of underestimating SDB severity without compromising specificity.
Abstract Study Objectives To quantify the amount of sleep stage ambiguity across expert scorers and to validate a new auto-scoring platform against sleep staging performed by multiple scorers. Methods We applied a new auto-scoring system to three datasets containing 95 PSGs scored by 6–12 scorers, to compare sleep stage probabilities (hypnodensity; i.e. the probability of each sleep stage being assigned to a given epoch) as the primary output, as well as a single sleep stage per epoch assigned by hierarchical majority rule. Results The percentage of epochs with 100% agreement across scorers was 46 ± 9%, 38 ± 10% and 32 ± 9% for the datasets with 6, 9, and 12 scorers, respectively. The mean intra-class correlation coefficient between sleep stage probabilities from auto- and manual-scoring was 0.91, representing excellent reliability. Within each dataset, agreement between auto-scoring and consensus manual-scoring was significantly higher than agreement between manual-scoring and consensus manual-scoring (0.78 vs. 0.69; 0.74 vs. 0.67; and 0.75 vs. 0.67; all p < 0.01). Conclusions Analysis of scoring performed by multiple scorers reveals that sleep stage ambiguity is the rule rather than the exception. Probabilities of the sleep stages determined by artificial intelligence auto-scoring provide an excellent estimate of this ambiguity. Compared to consensus manual-scoring, sleep staging derived from auto-scoring is for each individual PSG noninferior to manual-scoring meaning that auto-scoring output is ready for interpretation without the need for manual adjustment.
This study describes a computationally efficient algorithm for 4-class sleep staging based on cardiac activity and body movements. Using an accelerometer to calculate gross body movements and a reflective photoplethysmographic (PPG) sensor to determine interbeat intervals and a corresponding instantaneous heart rate signal, a neural network was trained to classify between wake, combined N1 and N2, N3 and REM sleep in epochs of 30 s. The classifier was validated on a hold-out set by comparing the output against manually scored sleep stages based on polysomnography (PSG). In addition, the execution time was compared with that of a previously developed heart rate variability (HRV) feature-based sleep staging algorithm. With a median epoch-per-epoch κ of 0.638 and accuracy of 77.8% the algorithm achieved an equivalent performance when compared to the previously developed HRV-based approach, but with a 50-times faster execution time. This shows how a neural network, without leveraging any a priori knowledge of the domain, can automatically “discover” a suitable mapping between cardiac activity and body movements, and sleep stages, even in patients with different sleep pathologies. In addition to the high performance, the reduced complexity of the algorithm makes practical implementation feasible, opening up new avenues in sleep diagnostics.
Human experts scoring sleep according to the American Academy of Sleep Medicine (AASM) rules are forced to select, for every 30-second epoch, one out of five stages, even if the characteristics of the neurological signals are ambiguous, a very common occurrence in clinical studies. Moreover, experts cannot score sleep in studies where these signals have not been recorded, such as in home sleep apnea testing (HSAT). In this topic review we describe how artificial intelligence can provide consistent and reliable scoring of sleep stages based on neurological signals recorded in polysomnography (PSG) and on cardiorespiratory signals recorded in HSAT. We also show how estimates of sleep stage probabilities, usually displayed as hypnodensity graph, can be used to quantify sleep stage ambiguity and stability. As an example of the application of hypnodensity in the characterization of sleep disordered breathing (SDB), we compared 49 patients with sleep apnea to healthy controls and revealed a severity-depending increase in ambiguity and decrease in stability during non-rapid eye movement (NREM) sleep. Moreover, using autoscoring of cardiorespiratory signals, we show how HSAT-derived apnea-hypopnea index and hypoxic burden are well correlated with the PSG indices in 80 patients, showing how using this technology can truly enable HSATs as alternatives to PSG to diagnose SDB.
Abstract Introduction In a recent analysis of PSG studies with six, nine and twelve human scorers, we showed that sleep stage ambiguity is the rule rather than the exception, and that sleep stage probabilities (SSPs) calculated with artificial intelligence provide an excellent estimate of this ambiguity. Here we investigate clinical and demographic factors driving this ambiguity and evaluate potential benefits of using hypnodensity in addition to the traditional hypnogram for sleep assessment. Methods SSPs were determined in 195 healthy subjects aged 20–95 years and in 49 apnea patients aged 29–73 years (AHI: 5.8–105.5 events/hour). In addition to the recommended sleep parameters, we derived the percentage of ambiguous sleep stages (SSP ≤ 0.95), the mean amount of ambiguity per sleep stage (1-SSP), sleep stage continuity (absolute difference in SSP between two epochs) and sleep stage stability (two adjacent epochs belonging to the same sleep stage with SSP > 0.95 each) as well as NREM sleep depth (weighted average of NREM SSPs). Results The percentages of ambiguous NREM epochs (±SD) increased significantly with age (Pearson’s r=0.72; p< 0.01) from 40±9% in young healthy subjects (20-39 years, n=61) to 50±9% in middle-aged healthy subjects (40-59 years, n=59) and to 62±11% in older healthy subjects (≥ 60 years, n=75). The percentages of ambiguous REM epochs did not change significantly across age (r=0.1). Apnea patients showed an increased NREM ambiguity (77±15% compared to 53±10% in age- and sex-matched healthy controls; p< 0.01), while REM ambiguity was only slightly increased (p< 0.05). Furthermore, sleep stage continuity, stability, and NREM sleep depth decrease significantly with age and are significantly lower in apnea patients than in healthy controls (p< 0.01). Conclusion Artificial intelligence-based autoscoring shows ambiguity in 40% of NREM epochs in young healthy subjects, increasing to 77% in apnea patients. Assigning a single sleep stage to each epoch and presenting sleep architecture as a traditional hypnogram may be misleading, especially for older subjects and patients with sleep-disordered breathing. A hypnodensity chart representing sleep stage probabilities reflects this sleep staging ambiguity, provides all the information contained in a hypnogram, and offers insights into sleep ambiguity, continuity, stability, and sleep depth. Support (if any)
Abstract Introduction There have been significant advances in machine learning in recent years. This means that powerful methods are now available for classification problems, such as scoring sleep stages from neurological or cardiorespiratory signals. In the present work, validation studies for both applications are presented. Methods To determine the 5 sleep stages from the neurological signals, 54 sleep-wake-related features were calculated and classified by a bidirectional long short-term memory (LSTM) network which had been trained on 1956 manual scorings of 588 PSGs from 294 subjects (supervised deep learning). To determine the 4 stages (wake, light sleep, deep sleep, REM) from cardiorespiratory signals, a convolutional neural network combined with LSTM layers was used for feature extraction and classification. This network had been trained on 685 PSGs from 391 subjects (Bakker et al. JCSM 2021). The networks obtained were validated in 428 PSGs with one and 10 PSGs with 12 manual scorings (neurological staging) as well as in 2 two datasets, each containing 296 ambulatory recordings (cardiorespiratory staging). Results Cohen’s kappa between autoscoring based on neurological signals and manual scoring was 0.74 (95%-confidence interval: 0.74-0.74) for the 428 PSGs. The intraclass correlation coefficient (ICC) for absolute agreement between autoscoring and manual scoring was for the AHI 0.97 (0.96-0.98), for the arousal index 0.79 (0.67-0.86) and for the PLMSI 0.91 (0.88-0.93). The ICC between the sleep stage probabilities derived from the 12 manual scorings and the artificial intelligence (AI) derived hypnodensity was 0.91 (0.91-0.91). Cohen’s kappa values for the cardiorespiratory sleep staging were 0.68 (0.68-0.68) and 0.64 (0.63-0.64) for the 2 datasets with 296 ambulatory recordings each. Conclusion All metrics from the PSG validation studies show substantial (Cohen’s kappa > 0.6) as well as good to excellent agreement (ICC > 0.75 or > 0.90) compared to manual scorings. As an added value of the AI-supported PSG evaluation, the probabilities of the sleep stages per epoch are determined (hypnodensity graph). The valid estimation of the sleep stages from cardiorespiratory signals by means of AI may result in improved clinical interpretation of home sleep apnea tests, which are increasingly used in the sleep-disordered breathing diagnostic pathway. Support (If Any)
Conventionally, sleep and associated events are scored visually by trained technologists according to the rules summarized in the American Academy of Sleep Medicine Manual. Since its first publication in 2007, the manual was continuously updated; the most recent version as of this writing was published in 2020. Human expert scoring is considered as gold standard, even though there is increasing evidence of limited interrater reliability between human scorers. Significant advances in machine learning have resulted in powerful methods for addressing complex classification problems such as automated scoring of sleep and associated events. Evidence is increasing that these autoscoring systems deliver performance comparable to manual scoring and offer several advantages to visual scoring: (1) avoidance of the rather expensive, time-consuming, and difficult visual scoring task that can be performed only by well-trained and experienced human scorers, (2) attainment of consistent scoring results, and (3) proposition of added value such as scoring in real time, sleep stage probabilities per epoch (hypnodensity), estimates of signal quality and sleep/wake-related features, identifications of periods with clinically relevant ambiguities (confidence trends), configurable sensitivity and rule settings, as well as cardiorespiratory sleep staging for home sleep apnea testing. This chapter describes the development of autoscoring systems since the first attempts in the 1970s up to the most recent solutions based on deep neural network approaches which achieve an accuracy that allows to use the autoscoring results directly for review and interpretation by a physician.
Abstract Introduction Scoring algorithms have the potential to increase polysomnography (PSG) scoring efficiency while also ensuring consistency and reproducibility. We sought to validate an updated sleep staging algorithm (Somnolyzer; Philips, Monroeville PA USA) against manual sleep staging, by analyzing a dataset we have previously used to report sleep staging variability across nine center-members of the Sleep Apnea Global Interdisciplinary Consortium (SAGIC). Methods Fifteen PSGs collected at a single sleep clinic were scored independently by technologists at nine SAGIC centers located in six countries, and auto-scored with the algorithm. Each 30-second epoch was staged manually according to American Academy of Sleep Medicine criteria. We calculated the intraclass correlation coefficient (ICC) and performed a Bland-Altman analysis comparing the average manual- and auto-scored total sleep time (TST) and time in each sleep stage (N1, N2, N3, rapid eye movement [REM]). We hypothesized that the values from auto-scoring would show good agreement and reliability when compared to the average across manual scorers. Results The participants contributing to the original dataset had a mean (SD) age of 47 (12) years and 80% were male. Auto-scoring showed substantial (ICC=0.60-0.80) or almost perfect (ICC=0.80-1.00) reliability compared to manual-scoring average, with ICCs (95% confidence interval) of 0.976 (0.931, 0.992) for TST, 0.681 (0.291, 0.879) for time in N1, 0.685 (0.299, 0.881) for time in N2, 0.922 (0.791, 0.973) for time in N3, and 0.930 (0.811, 0.976) for time in REM. Similarly, Bland-Altman analyses showed good agreement between methods, with a mean difference (limits of agreement) of only 1.2 (-19.7, 22.0) minutes for TST, 13.0 (-18.2, 44.1) minutes for N1, -13.8 (-65.7, 38.1) minutes for N2, -0.33 (-26.1, 25.5) minutes for N3, and -1.2 (-25.9, 23.5) minutes for REM. Conclusion Results support high reliability and good agreement between the auto-scoring algorithm and average human scoring for measurements of sleep durations. Auto-scoring slightly overestimated N1 and underestimated N2, but results for TST, N3 and REM were nearly identical on average. Thus, the auto-scoring algorithm is acceptable for sleep staging when compared against human scorers. Support (if any) Philips.
Abstract Introduction Scoring algorithms have the potential to increase polysomnography (PSG) scoring efficiency while also ensuring consistency and reproducibility. We sought to validate an updated event detection algorithm (Somnolyzer; Philips, Monroeville PA USA) against manual scoring, by analyzing a dataset we have previously used to report scoring variability across nine center-members of the Sleep Apnea Global Interdisciplinary Consortium (SAGIC). Methods Fifteen PSGs collected at a single sleep clinic were scored independently by technologists at nine SAGIC centers located in six countries, and auto-scored with the algorithm. Arousals, apneas, and hypopneas were identified according to the American Academy of Sleep Medicine recommended criteria. We calculated the intraclass correlation coefficient (ICC) and performed a Bland-Altman analysis comparing the average manual- and auto-scored apnea-hypopnea index (AHI), arousal index (ArI), apneas, obstructive apneas, central apneas, mixed apneas, and hypopneas. We hypothesized that the values from auto-scoring would show good agreement and reliability when compared to the average across manual scorers. Results Participants contributing to the original dataset had a mean (SD) age of 47 (12) years, AHI of 24.7 (18.2) events/hour, and 80% were male. The ICCs (95% confidence interval) between average manual- and auto-scoring were almost perfect (ICC=0.80–1.00) for AHI [0.989 (0.968, 0.996)], ArI [0.897 (0.729, 0.964)], hypopneas [0.992 (0.978, 0.997)], total apneas [0.973 (0.924, 0.991)], and obstructive apneas [0.919 (0.781, 0.972)], and moderately reliable (ICC=0.40–0.60] for central [0.537 (0.069, 0.815)] and mixed [0.502 (0.021, 0.798)] apneas. Similarly, Bland-Altman analyses supported good agreement for event detection between techniques, with a mean difference (limits of agreement) of only 1.45 (-3.22, 6.12) events/hour for AHI, total apneas 5.2 (-23.9, 34.3), obstructive apneas 1.8 (-45.9, 49.5), central apneas 1.8 (-9.7, 13.4), mixed apneas 1.6 (-14.8, 17.9), and hypopneas 4.3 (-12.4, 20.9). Conclusion Results support almost perfect reliability between auto-scoring and manual scoring of AHI, ArI, hypopneas, total apneas, and obstructive apneas, as well as moderate reliability for central and mixed apneas. There was good agreement between methods, with small mean differences; wider limits of agreement for specific type of apneas did not affect accuracy of the overall AHI. Thus, the auto-scoring algorithm appears reliable for event detection. Support (if any) Philips
Purpose: There is great interest in unobtrusive long-term sleep measurements using wearable devices based on reflective photoplethysmography (PPG). Unfortunately, consumer devices are not validated in patient populations and therefore not suitable for clinical use. Several sleep staging algorithms have been developed and validated based on ECG-signals. However, translation from these techniques to data derived by wearable PPG is not trivial, and requires the differences between sensing modalities to be integrated in the algorithm, or having the model trained directly with data obtained with the target sensor. Either way, validation of PPG-based sleep staging algorithms requires a large dataset containing both gold standard measurements and PPG-sensor in the applicable clinical population. Here, we take these important steps towards unobtrusive, long-term sleep monitoring. Methods: We developed and trained an algorithm based on wrist-worn PPG and accelerometry. The method was validated against reference polysomnography in an independent clinical population comprising 244 adults and 48 children (age: 3 to 82 years) with a wide variety of sleep disorders. Results: The classifier achieved substantial agreement on four-class sleep staging with an average Cohen's kappa of 0.62 and accuracy of 76.4%. For children/adolescents, it achieved even higher agreement with an average kappa of 0.66 and accuracy of 77.9%. Performance was significantly higher in non-REM parasomnias (kappa = 0.69, accuracy = 80.1%) and significantly lower in REM parasomnias (kappa = 0.55, accuracy = 72.3%). A weak correlation was found between age and kappa (rho = -0.30, p<0.001) and age and accuracy (rho = -0.22, p<0.001). Conclusion: This study shows the feasibility of automatic wearable sleep staging in patients with a broad variety of sleep disorders and a wide age range. Results demonstrate the potential for ambulatory long-term monitoring of clinical populations, which may improve diagnosis, estimation of severity and follow up in both sleep medicine and research.
Unobtrusive home sleep monitoring using wrist-worn wearable photoplethysmography (PPG) could open the way for better sleep disorder screening and health monitoring. However, PPG is rarely included in large sleep studies with gold-standard sleep annotation from polysomnography. Therefore, training data-intensive state-of-the-art deep neural networks is challenging. In this work a deep recurrent neural network is first trained using a large sleep data set with electrocardiogram (ECG) data (292 participants, 584 recordings) to perform 4-class sleep stage classification (wake, rapid-eye-movement, N1/N2, and N3). A small part of its weights is adapted to a smaller, newer PPG data set (60 healthy participants, 101 recordings) through three variations of transfer learning. Best results (Cohen's kappa of 0.65 ± 0.11, accuracy of 76.36 ± 7.57%) were achieved with the domain and decision combined transfer learning strategy, significantly outperforming the PPG-trained and ECG-trained baselines. This performance for PPG-based 4-class sleep stage classification is unprecedented in literature, bringing home sleep stage monitoring closer to clinical use. The work demonstrates the merit of transfer learning in developing reliable methods for new sensor technologies by reusing similar, older non-wearable data sets. Further study should evaluate our approach in patients with sleep disorders such as insomnia and sleep apnoea.
STUDY OBJECTIVES:We have developed the CardioRespiratory Sleep Staging (CReSS) algorithm for estimating sleep stages using heart rate variability and respiration, allowing for estimation of sleep staging during home sleep apnea tests. Our objective was to undertake an epoch-by-epoch validation of algorithm performance against the gold standard of manual polysomnography sleep staging. METHODS:Using 296 polysomnographs, we created a limited montage of airflow and heart rate and deployed CReSS to identify each 30-second epoch as wake, light sleep (N1 + N2), deep sleep (N3), or rapid eye movement (REM) sleep. We calculated Cohen's kappa and the percentage of accurately identified epochs. We repeated our analyses after stratification by sleep-disordered breathing (SDB) severity, and after adding thoracic respiratory effort as a backup signal for periods of invalid airflow. RESULTS:CReSS discriminated wake/light sleep/deep sleep/REM sleep with 78% accuracy; the kappa value was 0.643 (95% confidence interval, 0.641-0.645). Discrimination of wake/sleep demonstrated a kappa value of 0.711 and accuracy of 89%, non-REM sleep/REM sleep demonstrated a kappa of 0.790 and accuracy of 94%, and light sleep/deep sleep demonstrated a kappa of 0.469 and accuracy of 87%. Kappa values did not vary by more than 0.07 across subgroups of no SDB, mild SDB, moderate SDB, and severe SDB. Accuracy increased to 80%, with a kappa value of 0.680 (95% confidence interval, 0.678-0.682), when CReSS additionally utilized the thoracic respiratory effort signal. CONCLUSIONS:We observed substantial agreement between CReSS and the gold-standard comparator of manual sleep staging of polysomnographic signals, which was consistent across the full range of SDB severity. Future research should focus on the extent to which CReSS reduces the discrepancy between the apnea-hypopnea index and the respiratory event index, and the ability of CReSS to identify REM sleep-related obstructive sleep apnea.
Sleep and memory studies often focus on overnight rather than long-term memory changes, traditionally associating overnight memory change (OMC) with sleep architecture and sleep patterns such as spindles. In addition, (para-)sympathetic innervation has been associated with OMC after a daytime nap using heart rate variability (HRV). In this study we investigated overnight and long-term performance changes for procedural memory and evaluated associations with sleep architecture, spindle activity (SpA) and HRV measures (R-R interval [RRI], standard deviation of R-R intervals [SDNN], as well as spectral power for low [LF] and high frequencies [HF]). All participants (N = 20, M-age = 23.40 +/- 2.78 years) were trained on a mirror-tracing task and completed a control (normal vision) and learning (mirrored vision) condition. Performance was evaluated after training (R1), after a full-night sleep (R2) and 7 days thereafter (R3). Overnight changes (R2-R1) indicated significantly higher accuracy after sleep, whereas a significant long-term (R3-R2) improvement was only observed for tracing speed. Sleep architecture measures were not associated with OMC after correcting for multiple comparisons. However, individual SpA change from the control to the learning night indicated that only "SpA enhancers" exhibited overnight improvements for accuracy and long-term improvements for speed. HRV analyses revealed that lower SDNN and LF power was associated with better OMC for the procedural speed measure. Altogether, this study indicates that overnight improvement for procedural memory is specific for spindle enhancers, and is associated with HRV during sleep following procedural learning.
Study Objectives: To validate a previously developed sleep staging algorithm using heart rate variability (HRV) and body movements in an independent broad cohort of unselected sleep disordered patients. Methods: We applied a previously designed algorithm for automatic sleep staging using long short-term memory recurrent neural networks to model sleep architecture. The classifier uses 132 HRV features computed from electrocardiography and activity counts from accelerometry. We retrained our algorithm using two public datasets containing both healthy sleepers and sleep disordered patients. We then tested the performance of the algorithm on an independent hold-out validation set of sleep recordings from a wide range of sleep disorders collected in a tertiary sleep medicine center. Results: The classifier achieved substantial agreement on four-class sleep staging (wake/N1-N2/N3/rapid eye movement [REM)), with an average kappa of 0.60 and accuracy of 75.9%. The performance of the sleep staging algorithm was significantly higher in insomnia patients (kappa = 0.62, accuracy = 77.3%). Only in REM parasomnias, the performance was significantly lower (kappa = 0.47, accuracy = 70.5%). For two-class wake/sleep classification, the classifier achieved a kappa of 0.65, with a sensitivity (to wake) of 72.9% and specificity of 94.0%. Conclusions: This study shows that the combination of HRV, body movements, and a state-of-the-art deep neural network can reach substantial agreement in automatic sleep staging compared with polysomnography, even in patients suffering from a multitude of sleep disorders. The physiological signals required can be obtained in various ways, including non-obtrusive wrist-worn sensors, opening up new avenues for clinical diagnostics.
Bipolar disorder (BD) is a chronic illness with a relapsing and remitting time course. Relapses are manic or depressive in nature and intermitted by euthymic states. During euthymic states, patients lack the criteria for a manic or depressive diagnosis, but still suffer from impaired cognitive functioning as indicated by difficulties in executive and language-related processing. The present study investigated whether these deficits are reflected by altered intracortical activity in or functional connectivity between brain regions involved in these processes such as the prefrontal and the temporal cortices. Vigilance-controlled resting state EEG of 13 euthymic BD patients and 13 healthy age- and sex-matched controls was analyzed. Head-surface EEG was recomputed into intracortical current density values in 8 frequency bands using standardized low-resolution electromagnetic tomography. Intracortical current densities were averaged in 19 evenly distributed regions of interest (ROIs). Lagged coherences were computed between each pair of ROIs. Source activity and coherence measures between patients and controls were compared (paired t tests). Reductions in temporal cortex activity and in large-scale functional connectivity in patients compared to controls were observed. Activity reductions affected all 8 EEG frequency bands. Functional connectivity reductions affected the delta, theta, alpha-2, beta-2, and gamma band and involved but were not limited to prefrontal and temporal ROIs. The findings show reduced activation of the temporal cortex and reduced coordination between many brain regions in BD euthymia. These activation and connectivity changes may disturb the continuous frontotemporal information flow required for executive and language-related processing, which is impaired in euthymic BD patients.