With a focus on physical activity and physiological variables this scoping review synthesizes recent trends in machine learning for glycemic prediction in individuals with Type 1 diabetes. A structured PRISMA-ScR search (2010–2025) identified 41 studies which resulted in three dominant application areas: (1) Multi-horizon prediction of glycemia and physical activity detection, driven mainly by recurrent neural networks (RNN)-most commonly long short-term memory (LSTM)-with evidence that incorporating energy expenditure improves model performance; (2) prediction of exercise-induced dysglycemia and nocturnal hypoglycemia, which share overlapping temporal horizons, indicating potential for unified forecasting models; and (3) translation of prediction models into bolus-optimization strategies, though real-world validation is limited. The review identifies two critical gaps: (1) The handling of physiological drift and model decay as a result of physiological training or detraining; (2) Menstrual cycle integration and its use as a feature remains unexplored, while multiple studies have demonstrated the decrease of insulin sensitivity in the late luteal phase of the cycle.
Youth mental health-related problems and disorders have garnered increased attention due to global prevalence estimates that have, in some cases, increased following the COVID-19 pandemic. Various methodologies have been proposed to leverage artificial intelligence (AI) for detecting mental health problems in the general population; however, research specifically focused on AI methods for youth remains limited. Shortcomings in modern AI include limited training data modalities (i.e., types of input data used for model training), reliance on offline training, and the use of static models. This scoping review provides an overview of evidence that uses AI methods applied to youth mental health (YMH) and provides an assessment of the current state of research that integrates multimodal AI (i.e., models that incorporate multiple data modalities) and/or online learning (i.e., incremental or continual model training from streaming data) for the diagnosis, monitoring, and treatment of YMH-related problems. The findings indicate that research in AI applied to YMH is limited in the areas of multimodal AI and online learning. The number of studies in this field is steadily growing. Studies incorporating online learning demonstrate that this approach enhances model performance and adaptability, which is crucial for developing translational models capable of addressing real-world challenges effectively. Despite these advances, key challenges remain, including the availability and long-term validity of multimodal data, the lack of participant-related information in certain databases and studies, the ethical and logistical difficulties of collecting data from minors, and the computational costs of training robust AI models.
Background: There has been rapid growth in the field of deep learning with convolutional neural networks (CNNs) for imaging-related tasks. More recently, vision transformers (ViTs) have shown competitive performance to CNNs, while also uniquely possessing a novel 'attention mechanism'. ViTs may replace CNNs; however, more work is required to show if transformer architecture, specifically the new attention mechanism, is more receptive to salient information from input images. This issue is becoming increasingly critical as the use of machine learning continues to expand in healthcare. Thus, we proposed to assess whether the attention that ViTs receive is misplaced. Methods: The attention heads of ViT and the important pixels for a ResNet50 model were compared to radiologist annotations to determine appropriateness of each model's attention mechanisms in classifying two datasets. We used the VinDr-CXR dataset of 18,000 chest X-rays and the Mini-DDSM dataset of 10,000 mammograms to classify healthy, benign, and malignant tumours. The attention heads were examined through attention rollout and the ResNet50 model using the occlusion method. Models were evaluated with accuracy and level of agreement calculated with a pixel-wise logical XNOR operator. Results: The VinDr-CXR ResNet50 and the ViT models had test accuracy of 70.4% and 77.4%, respectively. The Mini-DDSM ResNet50 and ViT models had test accuracy of 90.26% and 95.53%, respectively. Agreement was higher for transformer models at 94.72% and 96.96% compared to 88.07% and 94.85% for the CNN models. Conclusions: We show that our attention is not misplaced as the transformer approach shows overall better performance and agreement compared to CNNs.
The C3PO collaborative, with a history of successful quality improvement (QI) initiatives, leveraged registry participants to develop a multi-center QI initiative to reduce adverse events (AEs) in congenital cardiac catheterization. A 32-person, interdisciplinary working group analyzed audited data for all congenital cardiac catheterization cases from 2014-2017. The primary outcome was the occurrence of any high-severity (level 3/4/5) AE. Cases were organized from shortest to longest duration, and level 3/4/5 and 4/5 AE rates were summarized for each procedure duration decile. Observations from the root cause analysis were used to inform the creation of a key driver diagram and determine change strategies and implementation tools. To facilitate pre-procedure communication and risk assessment, an online risk calculator was developed using 2014-2019 data. Between 2014-2017, 14,717 cases were entered from 10 sites. Level 3/4/5 AEs occurred in 732 (5.0%) cases, while 4/5 AEs occurred in 224 (1.5%) cases. The key driver diagram defined three drivers: (1) Pre-Procedure Risk Assessment, (2) Possibly Preventable Events, and (3) Procedure Length Optimization. Actionable change strategies organized around five communication timepoints were developed in interdisciplinary discussions. Pre-case risk calculator outputs were available as a case summary print out and incorporated into a calendar for weekly schedule planning. Pre-intervention (2019) and preliminary intervention period data (2020-2021) are presented here. Through improved resource planning, the protocol equips catheterization teams to respond efficiently to AEs and possibly prevent escalation into dangerous events. This protocol provides reproducible interventions that can be adapted to local practice.
Objective:The aim of this review is to identify gaps and provide a direction for future research in the utilization of Artificial Intelligence (AI) in chronic pain (CP) management.Methods:A comprehensive literature search was conducted using various databases, including Ovid MEDLINE, Web of Science Core Collection, IEEE Xplore, and ACM Digital Library. The search was limited to studies on AI in CP research, focusing on diagnosis, prognosis, clinical decision support, self-management, and rehabilitation. The studies were evaluated based on predefined inclusion criteria, including the reporting quality of AI algorithms used.Results:After the screening process, 60 studies were reviewed, highlighting AI’s effectiveness in diagnosing and classifying CP while revealing gaps in the attention given to treatment and rehabilitation. It was found that the most commonly used algorithms in CP research were support vector machines, logistic regression and random forest classifiers. The review also pointed out that attention to CP mechanisms is negligible despite being the most effective way to treat CP.Conclusion:The review concludes that to achieve more effective outcomes in CP management, future research should prioritize identifying CP mechanisms, CP management, and rehabilitation while leveraging a wider range of algorithms and architectures.Significance:This review highlights the potential of AI in improving the management of CP, which is a significant personal and economic burden affecting more than 30% of the world’s population. The identified gaps and future research directions provide valuable insights to researchers and practitioners in the field, with the potential to improve healthcare utilization.
Artificial intelligence (AI), specifically machine learning, has the potential to augment human decision-making in mental healthcare to improve clinical outcomes. An important challenge is the knowledge gap between AI designers and mental health professionals. We review AI researchers' and health professionals' related terminologies, concepts, and perspectives to facilitate interdisciplinary collaborations in AI-assisted decision-making research. We present a case study to understand the decision-making processes at McMaster Children's Hospital's Child and Youth Mental Health Outpatient Services Program. We identified three decision points that can be augmented using AI: diagnosis, prognosis, and assessment. We present a systematic literature review of forty publications using the PRISMA methodology to investigate AI in mental health. Most publications focused on developing classification and prediction models outside the clinical workflow. We report open research challenges investigating how health professionals' decision-making changes to AI, where AI can assist decision-making, and how the human–AI feedback loop informs model improvement.
Background The proportion of Canadian youth seeking mental health support from an emergency department (ED) has risen in recent years. As EDs typically address urgent mental health crises, revisiting an ED may represent unmet mental health needs. Accurate ED revisit prediction could aid early intervention and ensure efficient healthcare resource allocation. We examine the potential increased accuracy and performance of graph neural network (GNN) machine learning models compared to recurrent neural network (RNN), and baseline conventional machine learning and regression models for predicting ED revisit in electronic health record (EHR) data. Methods This study used EHR data for children and youth aged 4–17 seeking services at McMaster Children’s Hospital’s Child and Youth Mental Health Program outpatient service to develop and evaluate GNN and RNN models to predict whether a child/youth with an ED visit had an ED revisit within 30 days. GNN and RNN models were developed and compared against conventional baseline models. Model performance for GNN, RNN, XGBoost, decision tree and logistic regression models was evaluated using F1 scores. Results The GNN model outperformed the RNN model by an F1-score increase of 0.0511 and the best performing conventional machine learning model by an F1-score increase of 0.0470. Precision, recall, receiver operating characteristic (ROC) curves, and positive and negative predictive values showed that the GNN model performed the best, and the RNN model performed similarly to the XGBoost model. Performance increases were most noticeable for recall and negative predictive value than for precision and positive predictive value. Conclusions This study demonstrates the improved accuracy and potential utility of GNN models in predicting ED revisits among children and youth, although model performance may not be sufficient for clinical implementation. Given the improvements in recall and negative predictive value, GNN models should be further explored to develop algorithms that can inform clinical decision-making in ways that facilitate targeted interventions, optimize resource allocation, and improve outcomes for children and youth.
A 31-year-old woman with transposition of the great arteries status post-Senning operation presents with severe pulmonary venous baffle obstruction. Both standards of care (percutaneous stenting or open repair) were deemed suboptimal and/or high risk. A multidisciplinary, hybrid approach via subxiphoid incision, guided by 3-dimensional modeling, provided a lower risk and minimally invasive intervention.
This work-in-progress paper will examine the synergy of roles of instructional assistant interns (IAIs) and teaching assistants (TAs) in remote teaching and learning of the redesigned first-year engineering curriculum. Unique to this redesigned first-year engineering course for 1200 students in the current pandemic is the introduction of 11 IAIs on top of eight instructors, 150 TAs, and four staff. IAIs are upper-year engineering students and are responsible for the delivery of weekly Lab and Design Studio sessions while TAs are also upper-year engineering students who were hired on a part-time basis for students' mentorship and grading. Viewed through the lens of legitimate peripheral participation in a community of practice (Lave & Wenger, 1991), this paper will explore the best practices of combining IAIs and TAs in a course. A mixed methodology (Johnson & Onwuegbuzie, 2004; Purzer, 2011) will be employed, and data gathering and analysis are expected to be completed at the end of January 2021. A self-assessment survey regarding IAIs and TAs' roles, preparation, and experiences will be conducted. A separate focus group discussion and semi-structured interview with IAIs and TAs will also be undertaken. Quantitative data will be used to describe the reliability, spread, and correlations of responses between IAIs and TAs. Qualitative data will be coded thematically (Braun & Clarke, 2006). Themes will be generated using open-axial-selective coding (Corbin & Strauss, 2008). Quantitative and qualitative analysis will be triangulated to inform how IAIs and TAs promote performance skills crucial for future engineers such as communication abilities, conflict mediation, etc. (Bolstad et al., 2020). Best practices in teaching and learning through the synergies of IAIs and TAs, particularly in virtual setting, and lessons learned on how to address practical issues (like managing conflict), personal issues (like reflective practices), and professional development issues (like pedagogical training) will be highlighted.
Machine learning (ML) is a technique that learns to detect patterns and trends in data. However, the quality of reporting ML in research is often suboptimal, leading to inaccurate conclusions and hindering progress in the field, especially if disseminated in literature reviews that provide researchers with an overview of a field, current knowledge gaps, and future directions. While various tools are available to assess the quality and risk-of-bias of studies, there is currently no generalized tool for assessing the reporting quality of ML in the literature. To address this, this study presents a new screening tool called STAR-ML (Screening Tool for Assessing Reporting of Machine Learning), accompanied by a guide to using it. A pilot scoping review looking at ML in chronic pain was used to investigate the tool. The time it took to screen papers and how the selection of the threshold affected the papers included were explored. The tool provides researchers with a reliable and systematic way to evaluate the quality of reporting of ML studies and to make informed decisions about the inclusion of studies in scoping or systematic reviews. In addition, this study provides recommendations for authors on how to choose the threshold for inclusion and use the tool proficiently. Lastly, the STAR-ML tool can serve as a checklist for researchers seeking to develop or implement ML techniques effectively.
ObjectiveAs COVID-19 continues to affect the global population, it is crucial to study the impact of the disease in vulnerable populations. This study of a diverse, international cohort aims to provide timely, experiential data on the course of disease in paediatric patients with congenital heart disease (CHD). MethodsData were collected by capitalising on two pre-existing CHD registries, the International Quality Improvement Collaborative for Congenital Heart Disease: Improving Care in Low- and Middle-Income Countries and the Congenital Cardiac Catheterization Project on Outcomes. 35 participating sites reported data for all patients under 18 years of age with diagnosed CHD and known COVID-19 illness during 2020 identified at their institution. Patients were classified as low, moderate or high risk for moderate or severe COVID-19 illness based on patient anatomy, physiology and genetic syndrome using current published guidelines. Association of risk factors with hospitalisation and intensive care unit (ICU) level care were assessed. ResultsThe study included 339 COVID-19 cases in paediatric patients with CHD from 35 sites worldwide. Of these cases, 84 patients (25%) required hospitalisation, and 40 (12%) required ICU care. Age <1 year, recent cardiac intervention, anatomical complexity, clinical cardiac status and overall risk were all significantly associated with need for hospitalisation and ICU admission. A multivariable model for ICU admission including clinical cardiac status and recent cardiac intervention produced a c-statistic of 0.86. ConclusionsThese observational data suggest risk factors for hospitalisation related to COVID-19 in paediatric CHD include age, lower functional cardiac status and recent cardiac interventions. There is a need for further data to identify factors relevant to the care of patients with CHD who contract COVID-19 illness.
Medical events can affect space crew health and compromise the success of deep space missions. To successfully manage such events, crew members must be sufficiently prepared to manage certain medical conditions for which they are not technically trained. Extended Reality (XR) can provide an immersive, realistic user experience that, when integrated with augmented clinical tools (ACT), can improve training outcomes and provide real-time guidance during non-routine tasks, diagnostic, and therapeutic procedures. The goal of this study was to develop a framework to guide XR platform development using astronaut medical training and guidance as the domain for illustration. We conducted a mixed-methods study—using video conference meetings (45 subject-matter experts), Delphi panel surveys, and a web-based card sorting application—to develop a standard taxonomy of essential XR capabilities. We augmented this by identifying additional models and taxonomies from related fields. Together, this "taxonomy of taxonomies," and the essential XR capabilities identified, serve as an initial framework to structure the development of XR-based medical training and guidance for use during deep space exploration missions. We provide a schematic approach, illustrated with a use case, for how this framework and materials generated through this study might be employed.
Human Activity Recognition (HAR) has become a spotlight in recent scientific research because of its applications in various domains such as healthcare, athletic competitions, smart cities, and smart home. While researchers focus on the methodology of processing data, users wonder if the Artificial Intelligence (AI) methods used for HAR can be trusted. Trust depends mainly on the reliability or robustness of the system. To investigate the robustness of HAR systems, we analyzed several suitable current public datasets and selected WISDM for our investigation of Deep Learning approaches. While the published specification of WISDM matched our fundamental requirements (e.g., large, balanced, multi-hardware), several hidden issues were found in the course of our analysis. These issues reduce the performance and the overall trust of the classifier. By identifying the problems and repairing the dataset, the performance of the classifier was increased. This paper presents the methods by which other researchers may identify and correct similar problems in public datasets. By fixing the issues dataset veracity is improved, which increases the overall trust in the trained HAR system.
Despite their necessity in directing patient care worldwide, simple and accurate diagnostic tools for early Alzheimer’s disease (AD) do not exist. To support healthcare decision-making and planning, this research leverages large, multi-site accessible data and state-of-the-art supervised machine learning (XGBoost) to enable rapid, accurate, low-cost, accessible, non-invasive, interpretable, and early clinical evaluation of AD. Machine learning was employed to combine three key features: Everyday Cognition Questionnaire, Alzheimers Disease Assessment Scale, and Delayed Total Recall, achieving area under the receiver operating characteristic curves scores consistently above 97%. The selected features are important because they are non-invasive and easily collected. Low performance on delayed recall alone appears to distinguish most AD patients, consistent with the pathophysiology of AD where individuals having problems storing new information into long-term memory. Distinguishing this research from existing literature was the focus of enhancing the model's interpretability while maintaining performance of more complex and opaque models. The interpretable model enables understanding of the decision process, vital for clinical adoption of machine learning tools in AD evaluation. In summary, we present a methodology which identified accessible and noninvasive features, each with their absolute thresholds, together with a clinically operable decision route, to accurately and rapidly detect, differentiate, and diagnose Alzheimer's disease patients.
In this paper we present an augmented reality system for mobile devices that facilitates 3D brain tumor visualization in real time. The system uses facial features to track the subject in the scene. The system performs camera calibration based on the face size of the subject, instead of the common approach of using a number of chessboard images to calibrate the camera every time the application is installed on a new device. Camera 3D pose estimation is performed by finding its position and orientation based on a set of 3D points and their corresponding 2D projections. According to the estimated camera pose, a reconstructed brain tumor model is displayed at the same location as the subject’s real anatomy. The results of our experiment show the system was successful in performing the brain tumor augmentation in real time with a reprojection accuracy of 97%.
The increasing adoption of robot systems in industrial settings and teaming with humans have led to a growing interest in human-robot interaction (HRI) research. While many robots use sensors to avoid harming humans, they cannot elaborate on human actions or intentions, making them passive reactors rather than interactive collaborators. Intention-based systems can determine human motives and predict future movements, but their closer interaction with humans raises concerns about trust. This scoping review provides an overview of sensors, algorithms, and examines the trust aspect of intention-based systems in HRI scenarios. We searched MEDLINE, Embase, and IEEE Xplore databases to identify studies related to the forementioned topics of intention-based systems in HRI. Results from each study were summarized and categorized according to different intention types, representing various designs. The literature shows a range of sensors and algorithms used to identify intentions, each with their own advantages and disadvantages in different scenarios. However, trust of intention-based systems is not well studied. Although some research in AI and robotics can be applied to intention-based systems, their unique characteristics warrant further study to maximize collaboration performance. This review highlights the need for more research on the trust aspects of intention-based systems to better understand and optimize their role in human-robot interactions, at the same time establishes a foundation for future research in sensor and algorithm designs for intention-based systems.
Ultrasound (US) is the most widely used medical imaging modality due to its low cost, portability, real time imaging ability and use of non-ionizing radiation. However, unlike other imaging modalities such as CT or MRI, it is a heavily operator dependent, requiring trained expertise to leverage these benefits. Recently there has been an explosion of interest in AI across the medical community and many are turning to the growing trend of deep learning (DL) models to assist in diagnosis. However, due to possible differences in training and deployment, model performance suffers which can lead to misdiagnosis and operator hesitancy. This issue is known as dataset shift. Two aims to address dataset shift were proposed. The first was to quantify how US operator skill and hardware affects acquired images. The second was to use this skill quantification method to screen and match data to deep learning models to improve performance. A CAE Healthcare BLUE phantom with mock lesions was scanned by three operators using three different US systems (Siemens S3000, Clarius L15, and Ultrasonix SonixTouch) producing 39013 images. DL models were trained on a specific set to classify the presence of a simulated tumour and tested with data from differing sets. Principle Component Analysis (PCA) for dimension reduction was applied, then K-Means clustering was used to separate images generated by operator and hardware into clusters. This clustering algorithm was then used to screen incoming images during deployment to best match input to an appropriate DL model which is trained specifically to classify that type of operator or hardware. Results showed a noticeable difference when models were given data from differing datasets with the largest accuracy drop being 81.26% to 31.26%. Overall, operator differences more significantly affected DL model performance. Clustering models had much higher success separating hardware data compared to operator data. The proposed method reflects this result with a much higher accuracy across the hardware test set compared to the operator data.
Bayesian networks are increasingly used to quantify the uncertainty of subjective and stochastic concepts such as trust. In this article, we propose a data-driven approach to estimate Bayesian parameters in the domain of wearable medical devices. Our approach extracts the probability of a trust factor being in a specific state directly from the devices (e.g. sensor quality). The strength of the relationship between related factors is defined by expert knowledge and incorporated into the model. We use propagation rules from requirements engineering to estimate how much each trust factor contributes to the related intermediate nodes in the network and ultimately compute the trust score. The trust score is a relative measure of trustworthiness when different devices are evaluated in the same test conditions and using the same Bayesian structure. To evaluate our approach, we developed Bayesian networks for the trust quantification of similar wearable devices from two manufacturers under identical test conditions and noise levels. The results demonstrated the learnability and generalizability of our approach.
In this paper, we propose a data-driven approach to estimate Bayesian parameters when trust needs to be quantified in the domain of wearable medical devices (WMD). Our approach extracts the probability of a trust determinant (e.g., reliability or robustness) being in a specific state from the data. Then, we use the Bayesian approach to estimate the parameters for the intermediate nodes in the network and ultimately compute the trust score. The trust score we compute is used as a relative measure of trustworthiness between different WMDs evaluated in the same test conditions and with the same Bayesian network (BN). To evaluate our approach, we develop a BN for the trust quantification of similar wearable medical devices from two manufacturers under identical test conditions. The results demonstrate the learnability and generalizability of our data-driven parameter estimation approach.
The rapid integration of artificial intelligence across traditional research domains has generated an amalgamation of nomenclature. As cross-discipline teams work together on complex machine learning challenges, finding a consensus of basic definitions in the literature is a more fundamental problem. As a step in the Delphi process to define issues with trust and barriers to the adoption of autonomous systems, our study first collected and ranked the top concerns from a panel of international experts from the fields of engineering, computer science, medicine, aerospace, and defence, with experience working with artificial intelligence. This document presents a summary of the literature definitions for nomenclature derived from expert feedback.