Generative AI is reshaping healthcare, yet most existing advances rely on hospital-grade devices, which limits their accessibility and potential for health management outside clinical settings. With the proliferation of portable devices and telemedicine, healthcare is shifting toward home-based Diagnosis-It-Yourself (DIY) care. Despite this promise, several distinctive challenges remain: (i) home-collected data are heterogeneous, exacerbated by the absence of standardized large-scale datasets; (ii) models require adaptation to variable task demands and evolving individual conditions; (iii) the broad spectrum of home care tasks lacks a unified benchmark for systematic evaluation. In this paper, we present , a comprehensive framework designed to address these challenges through a tailored dataset, model, and benchmark. We first curate , a large-scale multimodal dataset capturing diverse real-world home care scenarios. Building on this, we propose , an adaptive foundation model for home-based health management, powered by the novel Hybrid Hyper Low-Rank Adaptation technique. Finally, we establish , the first benchmark to evaluate foundation models on home care tasks. Extensive experiments demonstrate that DIYHealthGPT delivers state-of-the-art performance over both general-purpose and medical-specific baselines on 11 home care tasks in both open-QA and closed-QA settings, laying the groundwork for the next generation of personalized health management at home.
Generative artificial intelligence (GenAI) is taking the world by storm. It promises transformative opportunities for advancing and disrupting existing practices, including healthcare. From large language models (LLMs) for clinical note synthesis and conversational assistance to multimodal systems that integrate medical imaging, electronic health records (EHRs), and genomic data for decision support, GenAI is transforming the practice of medicine and the delivery of healthcare, such as diagnosis and personalized treatments, with great potential in reducing the cognitive burden on clinicians, thereby improving overall healthcare delivery. However, GenAI deployment in healthcare requires an in-depth understanding of healthcare tasks and what can and cannot be achieved. In this paper, we propose a data-centric paradigm in the design and deployment of GenAI systems for healthcare. Specifically, we reposition the data lifecycle by making the medical data ecosystem the foundational substrate for generative healthcare systems. This ecosystem is designed to sustainably support the integration, representation, and retrieval of diverse medical data and knowledge. With effective and efficient data processing pipelines, such as semantic vector search and contextual querying, it enables GenAI-powered operations for upstream model components and downstream clinical applications. Ultimately, it not only supplies foundation models with high-quality, multimodal data for large-scale pretraining and domain-specific fine-tuning but also serves as a knowledge retrieval backend to support task-specific inference via the agentic layer. The ecosystem enables the deployment of GenAI for high-quality and effective healthcare delivery.
Transformer-based large language models (LLMs) have transformed the field of natural language processing and led to significant advancements in various text processing tasks. However, the applicability of these LLMs in identifying related drug-adverse event (AE) pairs within clinical context may be limited by the prevalent use of non-standard sentence structures and grammar. Nine transformer-based LLMs pre-trained on biomedical domain corpora are fine-tuned on annotated data (n = 5088) to classify drug-AE pairs in unstructured discharge summaries as causally related or unrelated. These LLMs are then validated on text segments from deidentified hospital discharge summaries from Singapore (n = 1647). To assess generalisability, the models are validated on annotated segments (n = 4418) from the Medical Information Mart for Intensive Care (MIMIC-III) database. Performance of LLMs in identifying related drug-AE pairs is then compared against a prior benchmark set by traditional machine learning models on the same data. Using an LLM-Bidirectional long short-term memory (LLM-BiLSTM) architecture, transformer-based LLMs improve F1 score as compared to prior benchmark with BioM-ELECTRA-Large-BiLSTM showing an average F1 score improvement of 16.1
In the metaverse the physical space and the virtual space co-exist, and interact simultaneously. While the physical space is virtually enhanced with information, the virtual space is continuously refreshed with real-time, real-world information. To allow users to process and manipulate information seamlessly between the real and digital spaces, novel technologies must be developed. These include smart interfaces, new augmented realities, and efficient data storage, management, and dissemination techniques. In this paper, we first discuss some promising co-space applications. These applications offer opportunities that neither of the spaces can realize on its own. Then, we further discuss several emerging technologies that empower the construction of metaverse. After that, we discuss comprehensively the data centric challenges. Finally, we discuss and envision what are likely to be required from the database and system perspectives.
Background Echocardiographic indexes of aortic stenosis may not comprehensively reflect disease morbidity. Plasma proteomic profiling may add prognostic value in these patients. Methods and Results Proximity extension assays (Olink) of 183 circulating cardiovascular and inflammatory proteins were performed in a prospective follow‐up study of 122 asymptomatic/minimally symptomatic patients (mean±SD age, 69.1±10.9 years; 61% men) with moderate to severe aortic stenosis and preserved left ventricular ejection fraction. Protein signatures of higher‐risk echocardiographic subgroups were determined. Associations of proteins with the primary composite outcome (heart failure hospitalization, progression to New York Heart Association class III‐IV, or all‐cause mortality) were evaluated using competing risk analyses, with aortic valve replacement being the competing risk. Network analysis unveiled mutually exclusive communities of proteins and echocardiographic parameters, connected only through NT‐proBNP (N‐terminal pro‐B‐type natriuretic peptide). Members of the tumor necrosis factor receptor superfamily (TNFRSF1A, TNFRSF1B, and TNFRSF14), and trefoil factor‐3 were major hub proteins among the circulating biomarkers. Left ventricular global longitudinal strain >−15% was associated with higher levels of proteins, primarily of inflammation and immune regulation, whereas aortic valve area <1 cm 2 , E/e’ >15, and left atrial reservoir strain <20% were associated with higher levels of NT‐proBNP. Of 14 proteins associated with the primary end point, phospholipase‐C, C‐X‐C motif chemokine‐9, and interleukin‐10 receptor subunit β demonstrated the highest hazard ratios after adjusting for clinical factors ( q <0.05). Conclusions Plasma proteins involved in inflammation and immune regulation were differentially expressed in patients with aortic stenosis with reduced left ventricular global longitudinal strain, and associated with adverse clinical outcomes. Their incorporation into aortic stenosis risk stratification warrants further assessment.
Localized food datasets have profound meaning in revealing a country’s special cuisines to explore people’s dietary behaviors, which will shed light on their health conditions and disease development. In this paper, revolving around the demand for accurate food recognition in Singapore, we develop the FoodSG platform to incu-bate diverse healthcare-oriented applications as a service in Singapore, taking into account their shared requirements. We release a localized Singaporean food dataset FoodSG-233 with a systematic cleaning and curation pipeline for promoting future data management research in food computing. To overcome the hurdle in recognition performance brought by Singaporean multifarious food dishes, we propose to integrate supervised contrastive learning into our food recognition model FoodSG-SCL for the intrinsic capability to mine hard positive/negative samples and therefore boost the accuracy. Through a comprehensive evaluation, we share the insightful experience with practitioners in the data management community regarding food-related data-intensive healthcare applications.
BACKGROUND Current cardiac magnetic resonance (CMR) imaging in pulmonary arterial hypertension (PAH) focuses on measures of ventricular function and coupling. OBJECTIVES The purpose of this study was to evaluate pulmonary artery (PA) global longitudinal strain (GLS) as a prognostic marker in patients with PAH. METHODS The authors included 169 patients with PAH from the ASPIRE (Assessing the Spectrum of Pulmonary hypertension Identified at a REferral centre) and INITIATE (Integrated computatioNal modelIng of righT heart mechanIcs and blood flow dynAmics in congeniTal hEart disease) registries, and 82 normal controls with similar age and gender distributions. PA GLS was derived from CMR feature tracking. Right ventricular measurements including volumes, ejection fraction, and right ventricular GLS were also derived from CMR. Patients were followed up a median of 34 months with all-cause mortality as the primary endpoint. Other known risk scores were collected, including the REVEAL (Registry to Evaluate Early and Long-term Pulmonary Arterial Hypertension Disease Management) 2.0 and COMPERA (Comparative, Prospective Registry of Newly Initiated Therapies for Pulmonary Hypertension) 2.0 scores. RESULTS Of 169 patients (mean age: 57 +/- 15 years; 80% female), 45 (26.6%) died (median follow-up: 34 months). Mean PA GLS was 23% +/- 6% in normal controls and 10% +/- 5% in patients with PAH (P< 0.0001). Patients with PA GLS<9% had a higher risk of mortality than those with PA GLS $9% (P < 0.001), and this was an independent predictor of mortality in PAH on multivariable analysis after adjustment for known risk factors (HR: 2.93; P = 0.010). Finally, in patients with PAH, PA GLS provided incremental prognostic value over the REVEAL 2.0 (global chi-square; P = 0.001; C statistic comparison; P = 0.030) and COMPERA 2.0 (global chi-square; P = 0.001; C statistic comparison; P = 0.048). CONCLUSIONS PA GLS confers incremental prognostic utility over the established risk scores for identifying patients with PAH at higher risk of death, who may be targeted for closer monitoring and/or intensified therapy. (J Am Coll Cardiol Img 2023;16:1022-1034) (c) 2023 by the American College of Cardiology Foundation.
In the Metaverse, the physical space and the virtual space co-exist, and interact simultaneously. While the physical space is virtually enhanced with information, the virtual space is continuously refreshed with real-time, real-world information. To allow users to process and manipulate information seamlessly between the real and digital spaces, novel technologies must be developed. These include smart interfaces, new augmented realities, efficient storage and data management and dissemination techniques. In this paper, we first discuss some promising co-space applications. These applications offer opportunities that neither of the spaces can realize on its own. We then discuss challenges. Finally, we discuss and envision what are likely to be required from the database and system perspectives.
Electronic Health Records (EHR) are generated from clinical routine care recording valuable information of broad patient populations, which provide plentiful opportunities for improving patient management and intervention strategies in clinical practice. To exploit the enormous potential of EHR data, a popular EHR data analysis paradigm in machine learning is EHR representation learning, which first leverages the individual patient's EHR data to learn informative representations by a backbone, and supports diverse health-care downstream tasks grounded on the representations. Unfortunately, such a paradigm fails to access the in-depth analysis of patients' relevance, which is generally known as cohort studies in clinical practice. Specifically, patients in the same cohort tend to share similar characteristics, implying their resemblance in medical conditions such as symptoms or diseases. In this paper, we propose a universal COhort Representation lEarning (CORE) framework to augment EHR utilization by leveraging the fine-grained cohort information among patients. In particular, CORE first develops an explicit patient modeling task based on the prior knowledge of patients' diagnosis codes, which measures the latent relevance among patients to adaptively divide the cohorts for each patient. Based on the constructed cohorts, CORE recodes the pre-extracted EHR data representation from intra- and inter-cohort perspectives, yielding augmented EHR data representation learning. CORE is readily applicable to diverse backbone models, serving as a universal plug-in framework to infuse cohort information into healthcare methods for boosted performance. We conduct an extensive experimental evaluation on two real-world datasets, and the experimental results demonstrate the effectiveness and generalizability of CORE.
Heart is the most important organ of the human body, and Electrocardiogram (ECG) is an essential tool for clinical monitoring of heart health and detecting cardiovascular diseases. Automatic detection of ECG anomalies is of great significance and clinical value in healthcare. However, performing automatic anomaly detection for the ECG data is challenging because we not only need to accurately detect the anomalies but also need to provide clinically meaningful interpretation of the results. Existing works on automatic ECG anomaly detection either rely on hand-crafted designs of feature extraction algorithms which are typically too simple to deliver good performance, or deep learning for automatically extracting features, which is not interpretable. In this paper, we propose ECGGAN, a novel reconstruction-based ECG anomaly detection framework. The key idea of ECGGAN is to make full use of the characteristics of ECG with the periodic metadata, namely beat, to learn the universal pattern in ECG from representative normal data. We establish a reconstruction model, taking leads as constraints to capture the unique characteristics of different leads in ECG data, and achieve accurate anomaly detection at ECG-level by combining multiple leads. Experimental results on two real-world datasets and their mixed-set confirm that our method achieves superior performance than baselines in terms of precision, recall, F1-score, and AUC. In addition, ECGGAN can provide clinically meaningful interpretation of results by revealing the extent to which abnormal sites deviate from the normal pattern.
Background: The role of left atrial (LA) strain as an imaging biomarker in aortic stenosis is not well estab-lished. The aim of this study was to investigate the prognostic performance of phasic LA strain in relation to clinical and echocardiographic variables and N-terminal pro-B-type natriuretic peptide in asymptomatic and minimally symptomatic patients with moderate to severe aortic stenosis and left ventricular ejection fraction > 50%.Methods: LA reservoir strain (LASr), LA conduit strain (LAScd), and LA contractile strain (LASct) were measured using speckle-tracking echocardiography. The primary outcome was a composite of all-cause mor-tality, heart failure hospitalization, progression to New York Heart Association functional class III or IV, acute coronary syndrome, or syncope. Secondary outcomes 1 and 2 comprised the same end points but excluded acute coronary syndrome and additionally syncope, respectively. The prognostic performance of phasic LA strain cutoffs was evaluated in competing risk analyses, aortic valve replacement being the competing risk.Results: Among 173 patients (mean age, 69 +/- 11 years; mean peak transaortic velocity, 4.0 +/- 0.8 m/sec), me-dian LASr, LAScd, and LASct were 27% (interquartile range [IQR], 22%-32%), 12% (IQR, 8%-15%), and 16% (IQR, 13%-18%), respectively. Over a median of 2.7 years (IQR, 1.4-4.6 years), the primary outcome and sec-ondary outcomes 1 and 2 occurred in 66 (38%), 62 (36%), and 59 (34%) patients, respectively. LASr < 20%, LAScd < 6%, and LASct < 12% were identified as optimal cutoffs of the primary outcome. In competing risk analyses, progressing from echocardiographic to echocardiographic-clinical and combined models incorpo-rating N-terminal pro-B-type natriuretic peptide, LA strain parameters outperformed other key echocardio-graphic variables and significantly predicted clinical outcomes. LASr < 20% was associated with the primary outcome and secondary outcome 1, LAScd < 6% with all clinical outcomes, and LASct < 12% with secondary outcome 2. LAScd < 6% had the highest specificity (95%) and positive predictive value (82%) for the primary outcome, and competing risk models incorporating LAScd < 6% had the best discriminative value.Conclusions: In well-compensated patients with moderate to severe aortic stenosis and preserved left ventric-ular ejection fractions, LA strain was superior to other echocardiographic indices and incremental to N -termi-nal pro-B-type natriuretic peptide for risk stratification. LAScd < 6%, LASr < 20%, and LASct < 12% identified patients at higher risk for adverse outcomes. (J Am Soc Echocardiogr 2023;36:29-37.)
The VGF gene encodes a neuronal secretory-peptide precursor that is rapidly induced by neurotrophic growth factors and by depolarization in vitro. VGF expression in the animal peaks during critical periods in the developing peripheral and central nervous systems. To gain insight into the possible functions and regulation of VGF in vivo, we have used in situ hybridization to examine the regulation of VGF messenger RNA by experimental manipulations, and have found it to be regulated in the CNS by paradigms that affect electrical activity and by lesion. Inhibition of retinal electrical activity during the critical period of visual development rapidly repressed VGF messenger RNA in the dorsal lateral geniculate nucleus of the thalamus. In the adult, kainate-induced seizures transiently induced VGF messenger RNA in neurons of the dentate gyrus, hippocampus, and cerebral cortex within hours. Cortical lesion strongly induced VGF messenger RNA in ipsilateral cortex within hours, and strongly repressed expression in ipsilateral striatum. Ten days postlesion there was a delayed induction of VGF messenger RNA in a portion of deafferented striatum where compensatory cortical sprouting has been detected. Expression of the neuronal secretory-peptide precursor VGF is therefore modulated in vivo by monocular deprivation, seizure, and cortical lesion, paradigms which lead to neurotrophin induction, synaptic remodeling and axonal sprouting.