Cerebral Palsy (CP) is a prevalent motor disability in children, for which early detection can significantly improve treatment outcomes. While skeleton-based Graph Convolutional Network (GCN) models have shown promise in automatically predicting CP risk from infant videos, their “black-box” nature raises concerns about clinical explainability. To address this, we introduce a perturbation framework tailored for infant movement features and use it to compare two explainable AI (XAI) methods: Class Activation Mapping (CAM) and Gradient-weighted Class Activation Mapping (Grad-CAM). First, we identify significant and non-significant body keypoints in very low and very high risk infant video snippets based on the XAI attribution scores. We then conduct targeted velocity and angular perturbations, both individually and in combination, on these keypoints to assess how the GCN model’s risk predictions change. Our results indicate that velocity-driven features of the arms, hips, and legs appear to have a dominant influence on CP risk predictions, while angular perturbations have a more modest impact. Furthermore, CAM and Grad-CAM show partial convergence in their explanations for both low and high CP risk groups. Our findings demonstrate the use of XAI-driven movement analysis for early CP prediction, and offer insights into potential movement-based biomarker discovery that warrant further clinical validation.
Early detection of Cerebral Palsy (CP) is crucial for effective intervention and monitoring. This paper tests the reliability and applicability of Explainable AI (XAI) methods using a deep learning method that predicts CP by analyzing skeletal data extracted from video recordings of infant movements. Specifically, we use XAI evaluation metrics — namely faithfulness and stability — to quantitatively assess the reliability of Class Activation Mapping (CAM) and Gradient-weighted Class Activation Mapping (Grad-CAM) in this specific medical application. We utilize a unique dataset of infant movements and apply skeleton data perturbations without distorting the original dynamics of the infant movements. Our CP prediction model utilizes an ensemble approach, so we evaluate the XAI metrics performances for both the overall ensemble and the individual models. Our findings indicate that both XAI methods effectively identify key body points influencing CP predictions and that the explanations are robust against minor data perturbations. Grad-CAM significantly outperforms CAM in the Relative Input Stability velocity (RISv) metric, which measures stability in terms of velocity. In contrast, CAM performs better in the Relative Input Stability bone (RISb) metric, which relates to bone stability, and the Relative Representation Stability (RRS) metric, which assesses internal representation robustness. Individual models within the ensemble show varied results, and neither CAM nor Grad-CAM consistently outperform the other, with the ensemble approach providing a representation of outcomes from its constituent models. Both CAM and Grad-CAM also perform significantly better than random attribution, supporting the robustness of these XAI methods. Ourwork demonstrates that XAI methods can offer reliable and stable explanations for CP prediction models. Future studies should further investigate how the explanations can enhance our understanding of specific movement patterns characterizing healthy and pathological development.
Explaining machine learning (ML) models using eXplainable AI (XAI) techniques has become essential to make them more transparent and trustworthy. This is especially important in high-risk environments like healthcare, where understanding model decisions is critical to ensure ethical, sound, and trustworthy outcome predictions. However, users are often confused about which explanability method to choose for their specific use case. We present a comparative analysis of two explainability methods, Shapley Additive Explanations (SHAP) and Gradient-weighted Class Activation Mapping (Grad-CAM), within the domain of human activity recognition (HAR) utilizing graph convolutional networks (GCNs). By evaluating these methods on skeleton-based input representation from two real-world datasets, including a healthcare-critical cerebral palsy (CP) case, this study provides vital insights into both approaches’ strengths, limitations, and differences, offering a roadmap for selecting the most appropriate explanation method based on specific models and applications. We qualitatively and quantitatively compare the two methods, focusing on feature importance ranking and model sensitivity through perturbation experiments. While SHAP provides detailed input feature attribution, Grad-CAM delivers faster, spatially oriented explanations, making both methods complementary depending on the application’s requirements. Given the importance of XAI in enhancing trust and transparency in ML models, particularly in sensitive environments like healthcare, our research demonstrates how SHAP and Grad-CAM could complement each other to provide model explanations.
IMPORTANCE Early identification of cerebral palsy (CP) is important for early intervention, yet expert-based assessments do not permit widespread use, and conventional machine learning alternatives lack validity. OBJECTIVE To develop and assess the external validity of a novel deep learning-based method to predict CP based on videos of infants' spontaneous movements at 9 to 18 weeks' corrected age. DESIGN, SETTING, AND PARTICIPANTS This prognostic study of a deep learning-based method to predict CP at a corrected age of 12 to 89 months involved 557 infants with a high risk of perinatal brain injury who were enrolled in previous studies conducted at 13 hospitals in Belgium, India, Norway, and the US between September 10, 2001, and October 25, 2018. Analysis was performed between February 11, 2020, and September 23, 2021. Included infants had available video recorded during the fidgety movement period from 9 to 18 weeks' corrected age, available classifications of fidgety movements ascertained by the general movement assessment (GMA) tool, and available data on CP status at 12 months' corrected age or older. A total of 418 infants (75.0%) were randomly assigned to the model development (training and internal validation) sample, and 139 (25.0%) were randomly assigned to the external validation sample (1 test set). EXPOSURE Video recording of spontaneous movements. MAIN OUTCOMES AND MEASURES The primary outcome was prediction of CP. Deep learning-based prediction of CP was performed automatically from a single video. Secondary outcomes included prediction of associated functional level and CP subtype. Sensitivity, specificity, positive and negative predictive values, and accuracy were assessed. RESULTS Among 557 infants (310 [55.7%] male), the median (IQR) corrected age was 12 (11-13) weeks at assessment, and 84 infants (15.1%) were diagnosed with CP at a mean (SD) age of 3.4 (1.7) years. Data on race and ethnicity were not reported because previous studies (from which the infant samples were derived) used different study protocols with inconsistent collection of these data. On external validation, the deep learning-based CP prediction method had sensitivity of 71.4% (95% CI, 47.8%-88.7%), specificity of 94.1%(95% CI, 88.2%-97.6%), positive predictive value of 68.2% (95% CI, 45.1%-86.1%), and negative predictive value of 94.9% (95% CI, 89.2%-98.1%). In comparison, the GMA tool had sensitivity of 70.0% (95% CI, 45.7%-88.1%), specificity of 88.7% (95% CI, 81.5%-93.8%), positive predictive value of 51.9% (95% CI, 32.0%-71.3%), and negative predictive value of 94.4% (95% CI, 88.3%-97.9%). The deep learning method achieved higher accuracy than the conventional machine learning method (90.6% [95% CI, 84.5%-94.9%] vs 72.7% [95% CI, 64.5%-79.9%]; P < .001), but no significant improvement in accuracy was observed compared with the GMA tool (85.9%; 95% CI, 78.9%-91.3%; P = .11). The deep learning prediction model had higher sensitivity among infants with nonambulatory CP (100%; 95% CI, 63.1%-100%) vs ambulatory CP (58.3%; 95% CI, 27.7%-84.8%; P = .02) and spastic bilateral CP (92.3%; 95% CI, 64.0%-99.8%) vs spastic unilateral CP (42.9%; 95% CI, 9.9%-81.6%; P < .001). CONCLUSIONS AND RELEVANCE In this prognostic study, a deep learning-based method for predicting CP at 9 to 18 weeks' corrected age had predictive accuracy on external validation, which suggests possible avenues for using deep learning-based software to provide objective early detection of CP in clinical settings.
OBJECTIVES:To determine whether videos taken by parents of their infants' spontaneous movements were in accordance with required standards in the In-Motion-App, and whether the videos could be remotely scored by a trained General Movement Assessment (GMA) observer. Additionally, to assess the feasibility of using home-based video recordings for automated tracking of spontaneous movements, and to examine parents' perceptions and experiences of taking videos in their homes.DESIGN:The study was a multi-centre prospective observational study.SETTING:Parents/families of high-risk infants in tertiary care follow-up programmes in Norway, Denmark and Belgium.METHODS:Parents/families were asked to video record their baby in accordance with the In-Motion standards which were based on published GMA criteria and criteria covering lighting and stability of smartphone. Videos were evaluated as GMA 'scorable' or 'non-scorable' based on predefined criteria. The accuracy of a 7-point body tracker software was compared with manually annotated body key points. Parents were surveyed about the In-Motion-App information and clarity.PARTICIPANTS:The sample comprised 86 parents/families of high-risk infants.RESULTS:The 86 parent/families returned 130 videos, and 121 (96%) of them were in accordance with the requirements for GMA assessment. The 7-point body tracker software detected more than 80% of body key point positions correctly. Most families found the instructions for filming their baby easy to follow, and more than 90% reported that they did not become more worried about their child's development through using the instructions.CONCLUSIONS:This study reveals that a short instructional video enabled parents to video record their infant's spontaneous movements in compliance with the standards required for remote GMA. Further, an accurate automated body point software detecting infant body landmarks in smartphone videos will facilitate clinical and research use soon. Home-based video recordings could be performed without worrying parents about their child's development.TRIALS REGISTRATION NUMBER:NCT03409978.
Gait parameters such as stride length, width, and period, as well as their respective variabilities, are widely used as indicators of mobility and walking function. Foot placement and its variability have thus been applied in areas such as aging, fall risk, spinal cord injury, diabetic neuropathy, and neurological conditions. But a drawback is that these measures are presently best obtained with specialized laboratory equipment such as motion capture systems and instrumented walkways, which may not be available in many clinics and certainly not during daily activities. One alternative is to fix inertial measurement units (IMUs) to the feet or body to gather motion data. However, few existing methods measure foot placement directly, due to drift associated with inertial data. We developed a method to measure stride-to-stride foot placement in unconstrained environments, and tested whether it can accurately quantify gait parameters over long walking distances. The method uses ground contact conditions to correct for drift, and state estimation algorithms to improve estimation of angular orientation. We tested the method with healthy adults walking over-ground, averaging 93 steps per trial, using a mobile motion capture system to provide reference data. We found IMU estimates of mean stride length and duration within 1% of motion capture, and standard deviations of length and width within 4% of motion capture. Step width cannot be directly estimated by IMUs, although lateral stride variability can. Inertial sensors measure walks over arbitrary distances, yielding estimates with good statistical confidence. Gait can thus be measured in a variety of environments, and even applied to long-term monitoring of everyday walking.
Assessment of spontaneous movements can predict the long-term developmental disorders in high-risk infants. In order to develop algorithms for automated prediction of later disorders, highly precise localization of segments and joints by infant pose estimation is required. Four types of convolutional neural networks were trained and evaluated on a novel infant pose dataset, covering the large variation in 1424 videos from a clinical international community. The localization performance of the networks was evaluated as the deviation between the estimated keypoint positions and human expert annotations. The computational efficiency was also assessed to determine the feasibility of the neural networks in clinical practice. The best performing neural network had a similar localization error to the inter-rater spread of human expert annotations, while still operating efficiently. Overall, the results of our study show that pose estimation of infant spontaneous movements has a great potential to support research initiatives on early detection of developmental disorders in children with perinatal brain injuries by quantifying infant movements from video recordings with human-level performance.
This study investigated the explanatory power of a sensor fusion of two complementary methods to explain performance and its underlying mechanisms in ski jumping. A differential Global Navigation Satellite System (dGNSS) and a markerless video-based pose estimation system (PosEst) were used to measure the kinematics and kinetics from the start of the in-run to the landing. The study had two aims; firstly, the agreement between the two methods was assessed using 16 jumps by athletes of national level from 5 m before the take-off to 20 m after, where the methods had spatial overlap. The comparison revealed a good agreement from 5 m after the take-off, within the uncertainty of the dGNSS (±0.05m). The second part of the study served as a proof of concept of the sensor fusion application, by showcasing the type of performance analysis the systems allows. Two ski jumps by the same ski jumper, with comparable external conditions, were chosen for the case study. The dGNSS was used to analyse the in-run and flight phase, while the PosEst system was used to analyse the take-off and the early flight phase. The proof-of-concept study showed that the methods are suitable to track the kinematic and kinetic characteristics that determine performance in ski jumping and their usability in both research and practice.
Single-person human pose estimation facilitates markerless movement analysis in sports, as well as in clinical applications. Still, state-of-the-art models for human pose estimation generally do not meet the requirements of real-life applications. The proliferation of deep learning techniques has resulted in the development of many advanced approaches. However, with the progresses in the field, more complex and inefficient models have also been introduced, which have caused tremendous increases in computational demands. To cope with these complexity and inefficiency challenges, we propose a novel convolutional neural network architecture, called EfficientPose, which exploits recently proposed EfficientNets in order to deliver efficient and scalable single-person pose estimation. EfficientPose is a family of models harnessing an effective multi-scale feature extractor and computationally efficient detection blocks using mobile inverted bottleneck convolutions, while at the same time ensuring that the precision of the pose configurations is still improved. Due to its low complexity and efficiency, EfficientPose enables real-world applications on edge devices by limiting the memory footprint and computational cost. The results from our experiments, using the challenging MPII single-person benchmark, show that the proposed EfficientPose models substantially outperform the widely-used OpenPose model both in terms of accuracy and computational efficiency. In particular, our top-performing model achieves state-of-the-art accuracy on single-person MPII, with low-complexity ConvNets.
Ovarian cancer, whether it occurs within or near the ovaries, involves different types of malignancies. While most ovarian tumors are benign, the malignant forms make up the most fatal gynecologic malignancies in the United States, as well as in other countries that regularly screen women for neoplasia of the cervix. The ovaries are responsible for maturation of follicles, usually between ages 11 and 50. They regulate maturation of eggs, ovulation, and the cyclical production of sex steroid hormones. Various ovarian cells coordinate these biologic functions. Each of the cell types has the potential to become neoplastic. Tumors may also metastasize to the ovaries from the breasts, colon, appendix, stomach, or pancreas. Bilateral ovarian masses from primary mucin-secreting gastrointestinal tumors are called Krukenberg tumors. This chapter focuses on tumors of epithelial origin, sex cord and stromal tumors, and germ cell tumors.
Assessment of spontaneous movements can predict the long-term developmental outcomes in high-risk infants. In order to develop algorithms for automated prediction of later function based on early motor repertoire, high-precision tracking of segments and joints are required. Four types of convolutional neural networks were investigated on a novel infant pose dataset, covering the large variation in 1 424 videos from a clinical international community. The precision level of the networks was evaluated as the deviation between the estimated keypoint positions and human expert annotations. The computational efficiency was also assessed to determine the feasibility of the neural networks in clinical practice. The study shows that the precision of the best performing infant motion tracker is similar to the inter-rater error of human experts, while still operating efficiently. In conclusion, the proposed tracking of infant movements can pave the way for early detection of motor disorders in children with perinatal brain injuries by quantifying infant movements from video recordings with human precision.