The purpose of this study was to: (1) compare the relative efficacy of different combinations of three behavioral intervention strategies (i.e., personalized reminders, financial incentives, and anchoring) for establishing physical activity habits using an mHealth app and (2) to examine the effects of these different combined interventions on intrinsic motivation for physical activity and daily walking habit strength. A four-arm randomized controlled trial was conducted in a sample of college students (N = 161) who had a self-reported personal wellness goal of increasing their physical activity. Receiving cue-contingent financial incentives (i.e., incentives conditional on performing physical activity within ± one hour of a prespecified physical activity cue) combined with anchoring resulted in the highest daily step counts and greatest odds of temporally consistent walking during both the four-week intervention and the full eight-week study period. Cue-contingent financial incentives were also more successful at increasing physical activity and maintaining these effects post-intervention than traditional non-cue-contingent incentives. There were no differences in intrinsic motivation or habit strength between study groups at any time point. Financial incentives, particularly cue-contingent incentives, can be effectively used to support the anchoring intervention strategy for establishing physical activity habits. Moreover, mHealth apps are a feasible method for delivering the combined intervention technique of financial incentives with anchoring.
Deep learning architectures have an extremely high-capacity for modeling complex data in a wide variety of domains. However, these architectures have been limited in their ability to support complex prediction problems using insurance claims data, such as readmission at 30 days, mainly due to data sparsity issue. Consequently, classical machine learning methods, especially those that embed domain knowledge in handcrafted features, are often on par with, and sometimes outperform, deep learning approaches. In this paper, we illustrate how the potential of deep learning can be achieved by blending domain knowledge within deep learning architectures to predict adverse events at hospital discharge, including readmissions. More specifically, we introduce a learning architecture that fuses a representation of patient data computed by a self-attention based recurrent neural network, with clinically relevant features. We conduct extensive experiments on a large claims dataset and show that the blended method outperforms the standard machine learning approaches.
Recent progress on few-shot learning largely relies on annotated data for meta-learning: base classes sampled from the same domain as the novel classes. However, in many applications, collecting data for meta-learning is infeasible or impossible. This leads to the cross-domain few-shot learning problem, where there is a large shift between base and novel class domains. While investigations of the cross-domain few-shot scenario exist, these works are limited to natural images that still contain a high degree of visual similarity. No work yet exists that examines few-shot learning across different imaging methods seen in real world scenarios, such as aerial and medical imaging. In this paper, we propose the Broader Study of Cross-Domain Few-Shot Learning (BSCD-FSL) benchmark, consisting of image data from a diverse assortment of image acquisition methods. This includes natural images, such as crop disease images, but additionally those that present with an increasing dissimilarity to natural images, such as satellite images, dermatology images, and radiology images. Extensive experiments on the proposed benchmark are performed to evaluate state-of-art meta-learning approaches, transfer learning approaches, and newer methods for cross-domain few-shot learning. The results demonstrate that state-of-art meta-learning methods are surprisingly outperformed by earlier meta-learning approaches, and all meta-learning methods underperform in relation to simple fine-tuning by 12.8% average accuracy. Performance gains previously observed with methods specialized for cross-domain few-shot learning vanish in this more challenging benchmark. Finally, accuracy of all methods tend to correlate with dataset similarity to natural images, verifying the value of the benchmark to better represent the diversity of data seen in practice and guiding future research.
Despite the large number of patients in Electronic Health Records (EHRs), the subset of usable data for modeling outcomes of specific phenotypes are often imbalanced and of modest size. This can be attributed to the uneven coverage of medical concepts in EHRs. We propose OMTL, an Ontology-driven Multi-Task Learning framework, that is designed to overcome such data limitations.The key contribution of our work is the effective use of knowledge from a predefined well-established medical relationship graph (ontology) to construct a novel deep learning network architecture that mirrors this ontology. This enables common representations to be shared across related phenotypes, and was found to improve the learning performance. The proposed OMTL naturally allows for multi-task learning of different phenotypes on distinct predictive tasks. These phenotypes are tied together by their semantic relationship according to the external medical ontology. Using the publicly available MIMIC-III database, we evaluate OMTL and demonstrate its efficacy on several real patient outcome predictions over state-of-the-art multi-task learning schemes. The results of evaluating the proposed approach on six experiments show improvement in the area under ROC curve by 9% and by 8% in the area under precision-recall curve.
Many institutions within the healthcare ecosystem are making significant investments in AI technologies to optimize their business operations at lower cost with improved patient outcomes. Despite the hype with AI, the full realization of this potential is seriously hindered by several systemic problems, including data privacy, security, bias, fairness, and explainability. In this paper, we propose a novel canonical architecture for the development of AI models in healthcare that addresses these challenges. This system enables the creation and management of AI predictive models throughout all the phases of their life cycle, including data ingestion, model building, and model promotion in production environments. This paper describes this architecture in detail, along with a qualitative evaluation of our experience of using it on real world problems.
Increased availability of electronic health records (EHR) has enabled researchers to study various medical questions. Cohort selection for the hypothesis under investigation is one of the main consideration for EHR analysis. For uncommon diseases, cohorts extracted from EHRs contain very limited number of records - hampering the robustness of any analysis. Data augmentation methods have been successfully applied in other domains to address this issue mainly using simulated records. In this paper, we present ODVICE, a data augmentation framework that leverages the medical concept ontology to systematically augment records using a novel ontologically guided Monte-Carlo graph spanning algorithm. The tool allows end users to specify a small set of interactive controls to control the augmentation process. We analyze the importance of ODVICE by conducting studies on MIMIC-III dataset for two learning tasks. Our results demonstrate the predictive performance of ODVICE augmented cohorts, showing ~30% improvement in area under the curve (AUC) over the non-augmented dataset and other data augmentation strategies.
We demonstrate the usage of our FoodKG [3], a food knowledge graph designed to assist in food recommendation. This resource, which brings together recipes, nutrition, food taxonomies, and links into existing ontologies, is used to power a cognitive agent that performs knowledge-base question answering, primarily to help improve peoples' diets by guiding them towards better foods. The system demonstration involves three categories of questions: simple queries for nutritional information, comparisons of nutrients between di erent foods, and constraintbased queries to nd recipes matching certain criteria.
The proliferation of recipes and other food information on the Web presents an opportunity for discovering and organizing diet-related knowledge into a knowledge graph. Currently, there are several ontologies related to food, but they are specialized in specific domains, e.g., from an agricultural, production, or specific health condition point-of-view. There is a lack of a unified knowledge graph that is oriented towards consumers who want to eat healthily, and who need an integrated food suggestion service that encompasses food and recipes that they encounter on a day-to-day basis, along with the provenance of the information they receive. Our resource contribution is a software toolkit that can be used to create a unified food knowledge graph that links the various silos related to food while preserving the provenance information. We describe the construction process of our knowledge graph, the plan for its maintenance, and how this knowledge graph has been utilized in several applications. These applications include a SPARQL-based service that lets a user determine what recipe to make based on ingredients at hand while taking constraints such as allergies into account, as well as a cognitive agent that can perform natural language question answering on the knowledge graph. Resource Website: https://foodkg.github.io
In recent years, there has been growing interest in the use of fitness trackers and smartphone applications for promoting physical activity. Many of these applications use accelerometers to estimate the level of activity that users engage in and provide visual reports of a user's step counts. When provided, most recommendations are limited to popular general health advice. In our study, we develop an approach for providing data-driven and personalized recommendations for intraday activity planning. We generate an hour-by-hour activity plan that is based on the user's probability of adhering to the plan. The user's probability of adherence to the plan is personalized, based on his/her past activity patterns and current activity target. Using this approach, we can tailor notifications (e.g., reminders, encouragement) to each user. We can also dynamically update the user's activity plan at mid-day, if his/her actual activity deviates sufficiently from the original plan. In this paper, we describe an implementation of our approach and report our technical findings with respect to identifying typical activity patterns from historical data, predicting whether an activity target will be achieved, and adapting an activity plan based on a user's actual performance throughout the day.
Person-generated health data (PGHD) generated by wearable devices and smartphone applications are growing rapidly. There is increasing effort to employ advanced analytical methods to generate insights from these data in order to help people change their lifestyle and improve their health. PGHD—such as step counts, exercise logs, nutritional diaries, and sleep records—are often incomplete, inaccurate, and collected over too short a duration. Insufficient user engagement with wearable and mobile technologies, as well as lack of sensor validation, standardization of data collection, transparency of data processing assumptions, and accessibility to relevant data from consumer-grade sensors, also negatively affects data quality. The literature on data quality for PGHD is sparse and fragmented, providing little guidance to data analysts on how to assess and prioritize data quality concerns. In this paper, we summarize our experiences as data analysts working with PGHD, outline some of the challenges in using PGHD for insight generation, and discuss some established methods for addressing these challenges. We review the literature on PGHD data quality, identify the major stakeholders in the PGHD ecosystem, and apply an established data quality framework to present the most relevant data quality challenges for each stakeholder.
Contact precautions are complex behavioral interventions. To better understand barriers to compliance, we conducted a prospective study that compared the time burden for health care workers caring for contact precautions patients versus other patients. We found that nurses spent significantly more time in the rooms of contact precautions patients. There was no significant change in physician timing. Future studies need to evaluate workflow changes so that barriers to contact precaution implementation can be fully understood and addressed.
We design and implement an intelligent advisor system for assisting individuals in changing their physical activity behavior. This system monitors an individual's physical activity and provides personalized, dynamic, pace-based recommendations to help an individual attain his/her activity goal. It adapts to a person's real-life constraints and generates recommendations using data from the individual's own activity history as well as data from a cohort of people similar to the individual. Our system relies on frequent predictions of the likelihood of goal attainment throughout the day, so that the current conditions and context are accounted for.
Background. Control of Clostridium difficile infection (CDI) is an increasingly difficult problem for health care institutions. There are commonly recommended strategies to combat CDI transmission, such as oral vancomycin for CDI treatment, increased hand hygiene with soap and water for health care workers, daily environmental disinfection of infected patient rooms, and contact isolation of diseased patients. However, the efficacy of these strategies, particularly for endemic CDI, has not been well studied. The objective of this research is to develop a valid, agent-based simulation model (ABM) to study C. difficile transmission and control in a midsized hospital. Methods. We develop an ABM of a midsized hospital with agents such as patients, health care workers, and visitors. We model the natural progression of CDI in a patient using a Markov chain and the transmission of CDI through agent and environmental interactions. We derive input parameters from aggregate patient data from the 2007–2010 Wisconsin Hospital Association and published medical literature. We define a calibration process, which we use to estimate transition probabilities of the Markov model by comparing simulation results to benchmark values found in published literature. Results. In a comparison of CDI control strategies implemented individually, routine bleach disinfection of CDI-positive patient rooms provides the largest reduction in nosocomial asymptomatic colonization (21.8%) and nosocomial CDIs (42.8%). Additionally, vancomycin treatment provides the largest reduction in relapse CDIs (41.9%), CDI-related mortalities (68.5%), and total patient length of stay (21.6%). Conclusion. We develop a generalized ABM for CDI control that can be customized and further expanded to specific institutions and/or scenarios. Additionally, we estimate transition probabilities for a Markov model of natural CDI progression in a patient through calibration.