Human action recognition from videos has gained substantial focus due to its wide applications in the field of video understanding. Most of the existing approaches extract human skeleton data from videos to encode actions because of the invariance nature of the skeleton information with respect to lightning con-ditions and background changes. Despite their success in achieving high recognition accuracy, methods based on limited body joints fail to capture the nuances of subtle body parts which are highly relevant for discriminating similar actions. In this paper, we overcome this limitation by presenting a holistic frame-work for combining spatial and motion features from the body, face, and hands to develop a novel data representation termed "Deep Actions Stamps (DeepActs)" for video-based action recognition. Compared to the skeleton sequences based on limited body joints, DeepActs encode more effective spatio-temporal features that provide robustness against pose estimation noises and improve action recognition accuracy. We also present "DeepActsNet", a deep learning based ensemble model which learns convolutional and structural features from Deep Action Stamps for highly accurate action recognition. Experiments on three challenging action recognition datasets (NTU60, NTU120, and SYSU) show that the proposed model pro-duces significant improvements in the action recognition accuracy with less computational cost compared to the state-of-the-art methods.(c) 2023 Elsevier Ltd. All rights reserved.
Background: Assistive automatic seizure detection can empower human annotators to shorten patient monitoring data review times. We present a proof-of-concept for a seizure detection system that is sensitive, automated, patient-specific, and tunable to maximise sensitivity while minimizing human annotation times. The system uses custom data preparation methods, deep learning analytics and electroencephalography (EEG) data. Methods: Scalp EEG data of 365 patients containing 171,745 s ictal and 2,185,864 s interictal samples obtained from clinical monitoring systems were analysed as part of a crowdsourced artificial intelligence (AI) challenge. Participants were tasked to develop an ictal/interictal classifier with high sensitivity and low false alarm rates. We built a challenge platform that prevented participants from downloading or directly accessing the data while allowing crowdsourced model development. Findings: The automatic detection system achieved tunable sensitivities between 75.00% and 91.60% allowing a reduction in the amount of raw EEG data to be reviewed by a human annotator by factors between 142x, and 22x respectively. The algorithm enables instantaneous reviewer-managed optimization of the balance between sensitivity and the amount of raw EEG data to be reviewed. Interpretation: This study demonstrates the utility of deep learning for patient-specific seizure detection in EEG data. Furthermore, deep learning in combination with a human reviewer can provide the basis for an assistive data labelling system lowering the time of manual review while maintaining human expert annotation performance. Funding: IBM employed all IBM Research authors. Temple University employed all Temple University authors. The Icahn School of Medicine at Mount Sinai employed Eren Ahsen. The corresponding authors Stefan Harrer and Gustavo Stolovitzky declare that they had full access to all the data in the study and that they had final responsibility for the decision to submit for publication.
This paper presents a novel deep learning enabled, video based analysis framework for assessing the Unified Parkinson’s Disease Rating Scale (UPDRS) that can be used in the clinic or at home. We report results from comparing the performance of the framework to that of trained clinicians on a population of 32 Parkinson’s disease (PD) patients. In-person clinical assessments by trained neurologists are used as the ground truth for training our framework and for comparing the performance. We find that the standard sit-to-stand activity can be used to evaluate the UPDRS sub-scores of bradykinesia (BRADY) and posture instability and gait disorders (PIGD). For BRADY we find Fl-scores of 0.75 using our framework compared to 0.50 for the video based rater clinicians, while for PIGD we find 0.78 for the framework and 0.45 for the video based rater clinicians. We believe our proposed framework has potential to provide clinically acceptable end points of PD in greater granularity without imposing burdens on patients and clinicians, which empowers a variety of use cases such as passive tracking of PD progression in spaces such as nursing homes, in-home self-assessment, and enhanced tele-medicine.
Objective Tracking seizures is crucial for epilepsy monitoring and treatment evaluation. Current epilepsy care relies on caretaker seizure diaries, but clinical seizure monitoring may miss seizures. Wearable devices may be better tolerated and more suitable for long-term ambulatory monitoring. This study evaluates the seizure detection performance of custom-developed machine learning (ML) algorithms across a broad spectrum of epileptic seizures utilizing wrist- and ankle-worn multisignal biosensors. Methods We enrolled patients admitted to the epilepsy monitoring unit and asked them to wear a wearable sensor on either their wrists or ankles. The sensor recorded body temperature, electrodermal activity, accelerometry (ACC), and photoplethysmography, which provides blood volume pulse (BVP). We used electroencephalographic seizure onset and offset as determined by a board-certified epileptologist as a standard comparison. We trained and validated ML for two different algorithms: Algorithm 1, ML methods for developing seizure type-specific detection models for nine individual seizure types; and Algorithm 2, ML methods for building general seizure type-agnostic detection, lumping together all seizure types. Results We included 94 patients (57.4% female, median age = 9.9 years) and 548 epileptic seizures (11 066 h of sensor data) for a total of 930 seizures and nine seizure types. Algorithm 1 detected eight of nine seizure types better than chance (area under the receiver operating characteristic curve [AUC-ROC] = .648-.976). Algorithm 2 detected all nine seizure types better than chance (AUC-ROC = .642-.995); a fusion of ACC and BVP modalities achieved the best AUC-ROC (.752) when combining all seizure types together. Significance Automatic seizure detection using ML from multimodal wearable sensor data is feasible across a broad spectrum of epileptic seizures. Preliminary results show better than chance seizure detection. The next steps include validation of our results in larger datasets, evaluation of the detection utility tool for additional clinical seizure types, and integration of additional clinical information.
Epilepsy is a common neurological disorder characterized by recurrent epileptic seizures. These seizures have different intensities and might lead to accidents or, in the worst case, to sudden death. Therefore, being able to predict epileptic seizures would allow patients to be prepared, reducing the risk of injury. This paper focuses on epileptic seizure prediction using EEG (Electroencephalogram) signals. In contrast to the standard approach where the preictal state is assumed to have a constant duration in all the seizures of a patient, we propose a new method that labels each seizure individually exploiting clustering. Our labeling approach, which was applicable for 38% of the selected seizures, results in substantial improvements compared to the standard one. In fact, it reduces noise in the labels and improves the performance of the binary classifier used to distinguish the interictal and preictal states. Hence, our results suggest that the preictal duration is seizure-specific, not only patient-specific. Finally, we show that our method is able to predict 17 out of 18 (94%) seizures between 15 and 85 minutes, before seizure onset.
Accurate classification of seizure types plays a crucial role in the treatment and disease management of epileptic patients. Epileptic seizure types not only impact the choice of drugs but also the range of activities a patient can safely engage in. With recent advances being made towards artificial intelligence enabled automatic seizure detection, the next frontier is the automatic classification of seizure types. On that note, in this paper, we explore the application of machine learning algorithms for multiclass seizure type classification. We used the recently released TUH EEG seizure corpus (v1.4.0 and v1.5.2) and conducted a thorough search space exploration to evaluate the performance of a combination of various preprocessing techniques, machine learning algorithms, and corresponding hyperparameters on this task. We show that our algorithms can reach a weighted F1 score of up to 0.901 for seizure-wise cross validation and 0.561 for patient-wise cross validation thereby setting a benchmark for scalp EEG based multi-class seizure type classification.
Ensemble models comprising of deep Convolutional Neural Networks (CNN) have shown significant improvements in model generalization but at the cost of large computation and memory requirements. In this paper, we present a framework for learning compact CNN models with improved classification performance and model generalization. For this, we propose a CNN architecture of a compact student model with parallel branches which are trained using ground truth labels and information from high capacity teacher networks in an ensemble learning fashion. Our framework provides two main benefits: i) Distilling knowledge from different teachers into the student network promotes heterogeneity in feature learning at different branches of the student network and enables the network to learn diverse solutions to the target problem. ii) Coupling the branches of the student network through ensembling encourages collaboration and improves the quality of the final predictions by reducing variance in the network outputs. Experiments on the well established CIFAR-10 and CIFAR-100 datasets show that our Ensemble Knowledge Distillation (EKD) improves classification accuracy and model generalization especially in situations with limited training data. Experiments also show that our EKD based compact networks outperform in terms of mean accuracy on the test datasets compared to state-of-the-art knowledge distillation based methods.
Existing action recognition methods mainly focus on joint and bone information in human body skeleton data due to its robustness to complex backgrounds and dynamic characteristics of the environments. In this paper, we combine body skeleton data with spatial and motion features from face and two hands, and present "Deep Action Stamps (DeepActs)", a novel data representation to encode actions from video sequences. We also present "DeepActsNet", a deep learning based ensemble model which learns convolutional and structural features from Deep Action Stamps for highly accurate action recognition. Experiments on three challenging action recognition datasets (NTU60, NTU120, and SYSU) show that the proposed model trained using Deep Action Stamps produce considerable improvements in the action recognition accuracy with less computational cost compared to the state-of-the-art methods.
Falling can have fatal consequences for elderly people especially if the fallen person is unable to call for help due to loss of consciousness or any injury. Automatic fall detection systems can assist through prompt fall alarms and by minimizing the fear of falling when living independently at home. Existing vision-based fall detection systems lack generalization to unseen environments due to challenges such as variations in physical appearances, different camera viewpoints, occlusions, and background clutter. In this paper, we explore ways to overcome the above challenges and present Single Shot Human Fall Detector (SSHFD), a deep learning based framework for automatic fall detection from a single image. This is achieved through two key innovations. First, we present a human pose based fall representation which is invariant to appearance characteristics. Second, we present neural network models for 3d pose estimation and fall recognition which are resilient to missing joints due to occluded body parts. Experiments on public fall datasets show that our framework successfully transfers knowledge of 3d pose estimation and fall recognition learnt purely from synthetic data to unseen real-world data, showcasing its generalization capability for accurate fall detection in real-world scenarios.
Background Assistive automatic seizure detection can empower human annotators to shorten patient monitoring data review times. We present a proof-of-concept for Findings The automatic detection system achieved tunable sensitivities between 75.00% and 91.60% allowing to reduce the amount of raw EEG data to be reviewed by a human annotator by factors between 142x, and 22x respectively. The algorithm enables instantaneous reviewer-managed optimisation of the balance between sensitivity and the amount of raw EEG data to be reviewed. Interpretation This study demonstrates the utility of deep learning for patient-specific seizure detection in EEG data. Furthermore, deep learning in combination with a human reviewer can provide the basis for an assistive data labelling system lowering the time of manual review while maintaining human expert annotation performance. the data in the study that they final for the to for publication.
Falling can have fatal consequences for the elderly people especially if the fallen person is unable to call for help due to loss of consciousness or any other associated injury. Automatic fall detection systems can assist in overcoming this issue through prompt fall alarms which then allow the triggering of a third party response, and to minimize the fear of falling when living independently at home. Vision-based fall detection systems detect human regions in the scene and use information from these regions to train classifiers for fall recognition. However, the performance of these systems lack generalization to unseen environments due to factors such as errors in the human detection stage and the unavailability of large-scale fall datasets to learn robust features for fall recognition. In this paper, we present a deep learning based framework towards automatic fall detection from RGB images captured by a single camera. Our framework learns human skeleton and segmentation based fall representations purely from synthetic data generated in a virtual environment. This de-identifies personal information contained in the original images and preserves privacy which is highly desirable in health informatics. Experiments on challenging real-world fall datasets show that our framework performs successful transfer of fall recognition knowledge from synthetic to real-world data and achieves high sensitivity and specificity scores showcasing its generalization capability for highly accurate fall detection in unseen real-world environments.
This paper presents Densely Supervised Grasp Detector (DSGD), a deep learning framework which combines CNN structures with layer-wise feature fusion and produces grasps and their confidence scores at different levels of the image hierarchy (i.e., global-, region-, and pixel-levels). Specifically, at the global-level, DSGD uses the entire image information to predict a grasp. At the region-level, DSGD uses a region proposal network to identify salient regions in the image and uses a grasp prediction network to generate segmentations and their corresponding grasp poses of the salient regions. At the pixel-level, DSGD uses a fully convolutional network and predicts a grasp and its confidence at every pixel. During inference, DSGD selects the most confident grasp as the output. This selection from hierarchically generated grasp candidates overcomes limitations of the individual models. DSGD outperforms state-of-the-art methods on the Cornell grasp dataset in terms of grasp accuracy. Evaluation on a multi-object dataset and real-world robotic grasping experiments show that DSGD produces highly stable grasps on a set of unseen objects in new environments. It achieves 97% grasp detection accuracy and 90% robotic grasping success rate with real-time inference speed.
Recent research on grasp detection has focused on improving accuracy through deep CNN models, but at the cost of large memory and computational resources. In this paper, we propose an efficient CNN architecture which produces high grasp detection accuracy in real-time while maintaining a compact model design. To achieve this, we introduce a CNN architecture termed GraspNet which has two main branches: i) An encoder branch which downsamples an input image using our novel Dilated Dense Fire (DDF) modules - squeeze and dilated convolutions with dense residual connections. ii) A decoder branch which upsamples the output of the encoder branch to the original image size using deconvolutions and fuse connections. We evaluated GraspNet for grasp detection using offline datasets and a real-world robotic grasping setup. In experiments, we show that GraspNet achieves superior grasp detection accuracy compared to the stateof-the-art computation-efficient CNN models with real-time inference speed on embedded GPU hardware (Nvidia Jetson TX1), making it suitable for low-powered devices.
Grasp detection from visual data is a recognition problem, where the goal is to determine regions in images which correspond to high grasp-ability with respect to certain quality metric. Existing deep learning based approaches mainly focus on predicting grasps, where the quality of the predictions is largely influenced by the choice of the CNN architecture and the objective function used for learning grasp representations. This paper presents a deep learning framework termed EnsembleNet which learns to produce and evaluate grasps within a unified framework. To achieve this, we formulate grasp detection as a two step procedure: i) Grasp generation where, EnsembleNet generates four different grasp representations (regression grasp, joint regression-classification grasp, segmentation grasp, and heuristic grasp), and ii) Grasp evaluation where EnsembleNet produces confidence scores for the generated grasps and selects the grasp with the highest score as the output. We evaluated EnsembleNet for grasp detection on RGBD object datasets. The experiments show that the grasps produced by EnsembleNet are more accurate compared to the independent CNN models and the state-of-the-art grasp detection methods.
Motor imagery (MI) based Brain-Computer Interfaces (BCIs) are a viable option for giving locked-in syndrome patients independence and communicability. BCIs comprising expensive medical-grade EEG systems evaluated in carefully-controlled, artificial environments are impractical for take-home use. Previous studies evaluated low-cost systems; however, performance was suboptimal or inconclusive. Here we evaluated a low-cost EEG system, OpenBCI, in a natural environment and leveraged neurofeedback, deep learning, and wider temporal windows to improve performance. $\mu-$rhythm data collected over the sensorimotor cortex from healthy participants performing relaxation and right-handed MI tasks were used to train a multi-layer perceptron binary classifier using deep learning. We showed that our method outperforms previous OpenBCI MI-based BCIs, thereby extending the BCI capabilities of this low-cost device.
This paper presents an efficient framework to perform recognition and grasp detection of objects from RGB-D images of real scenes. The framework uses a novel architecture of hierarchical cascaded forests, in which object-class and grasp-pose probabilities are computed at different levels of an image hierarchy (e.g., patch and object levels) and fused to infer the class and the grasp of unseen objects. We introduce a novel training objective function that minimizes the uncertainties of the class labels and the grasp ground truths at the leaves of the forests, thereby enabling the framework to perform the recognition and grasp detection of objects. Our objective function is learned from features that are extracted from RGB-D point clouds of the objects. For that, we propose a novel method to encode an RGB-D point cloud into a representation that facilitates the use of large convolution neural networks to extract discriminative features from RGB-D images. We evaluate our framework on challenging object datasets, where we demonstrate that our framework outperforms the state-of-the-art methods in terms of object-recognition and grasp-detection accuracies. We also show experiments by using live video streams from a Kinect mounted on our in-house robotic platform.
This index covers all technical items—papers, correspondence, reviews, etc.—that appeared in this periodical during 2017, and items from previous years that were commented upon or corrected in 2017. Departments and other items may also be covered if they have been judged to have archival value. The Author Index contains the primary entry for each item, listed under the first author’s name. The primary entry includes the coauthors’ names, the title of the paper or other item, and its location, specified by the publication abbreviation, year, month, and inclusive pagination. The Subject Index contains entries describing the item under all appropriate subject headings, plus the first author’s name, the publication abbreviation, month, and year, and inclusive pages. Note that the item title is found only under the primary entry in the Author Index.
While deep convolutional neural networks have shown a remarkable success in image classification, the problems of inter-class similarities, intra-class variances, the effective combination of multi-modal data, and the spatial variability in images of objects remain to be major challenges. To address these problems, this paper proposes a novel framework to learn a discriminative and spatially invariant classification model for object and indoor scene recognition using multi-modal RGB-D imagery. This is achieved through three postulates: 1) spatial invariance $-$ this is achieved by combining a spatial transformer network with a deep convolutional neural network to learn features which are invariant to spatial translations, rotations, and scale changes, 2) high discriminative capability $-$ this is achieved by introducing Fisher encoding within the CNN architecture to learn features which have small inter-class similarities and large intra-class compactness, and 3) multi-modal hierarchical fusion$-$ this is achieved through the regularization of semantic segmentation to a multi-modal CNN architecture, where class probabilities are estimated at different hierarchical levels (i.e., image- and pixel-levels), and fused into a Conditional Random Field (CRF)-based inference hypothesis, the optimization of which produces consistent class labels in RGB-D images. Extensive experimental evaluations on RGB-D object and scene datasets, and live video streams (acquired from Kinect) show that our framework produces superior object and scene classification results compared to the state-of-the-art methods.