
Coded clinical data are crucial in biomedical informatics research. While it is well known that electronic medical records often contain coding errors, numerous studies rely on International Classification of Diseases (ICD) codes for phenotyping in cohort assembly, statistical analysis, and AI modeling. Although fairness hasbecome an important focus in AI research, the potential biases embedded in coded clinical data have received less attention. In this study, we employed a race- and sex-agnostic AI phenotyping model to assess coding fairness across 203 ICD code blocks within the Veterans Health Administration Clinical Data Warehouse. Our findings revealed variability in coding consistency across demographic subgroups, including sex, race, and ethnicity. Notably, over 50% of the code blocks exhibitedstatisticallysignificant differences in discrepancies between AI-generated and ICD-based phenotypesacross these demographic groups. These results suggest the need to recognize and address demographic-related coding discrepancies to ensure coding fairness.
Lung cancer remains a significant challenge in public health, ranking among the leading causes of cancer-related mortality. Low-dose computed tomography (LDCT) -based lung cancer screening has emerged as an effective tool for early detection, particularly in high-risk populations. However, interpreting lung nodule characteristics from radiology reports can often be time-consuming and labor-intensive due to the length and inherent ambiguity of the reports, even with standardized reporting requirements like Lung-RADS. Generating Lung-RADS assessments from original radiology reports is a significant task for radiologists. This study addresses these challenges by developing an in-context learning framework utilizing large language models (LLMs). In this process, we aimed to identify the best approach that accurately categorizes lung nodules and streamlines management decisions, providing robust and interpretable decision support. Overall, this research aims to reduce the time and effort of the radiologist in lung cancer screening, ultimately enhancing efficiency and accuracy and enabling timely and precise interventions.
Intimate partner homicide (IPH) remains a major yet understudied cause of maternal mortality among U.S. women of childbearing age (WCBA). We leveraged the National Violent Death Reporting System (NVDRS) and county-level Maternal Vulnerability Index (MVI) data from 2018-2022 to train three machine learning models-logistic regression, random forest, and XGBoost-to classify whether homicides were IPH. Among 11,498 homicides involving WCBA, 33% were IPH. XGBoost achieved the best performance (F1-score = 0.83, AUPRC = 0.87), prompting further examination of key predictors via model explainability. Results indicated that acute interpersonal conflicts (e.g., arguments, jealousy), prior IPV victimization, and structural vulnerabilities (e.g., reproductive healthcare access, physical environment) were influential predictors of IPH. By illustrating the interplay of individual, interpersonal, and broader community-level risk factors, our study shows how machine learning can inform multilevel strategies to prevent IPH and improve maternal health.
Documenting clients, screenings and vaccinations administered is of particular importance during mass vaccination, since information regarding uptake is critical for monitoring adverse effects and vaccine efficacy. This is especially essential when a newly-developed vaccine is being dispensed or when multiple doses of vaccine are needed per person. Despite these needs, there is no uniform or integrated system for effective vaccine data collection. This work focuses on modernizing public health infrastructure through informatics. In this paper, we describe and analyze five types of electronic technologies for registration and screening in vaccination clinics. We contrast their functionalities, usability and operations performance based on time-motion studies and service data collected during actual influenza vaccination campaigns. We evaluate their dispensing performance under an optimal dispensing clinic design. Our analysis shows that each of these electronic technologies can improve overall throughput by 16% to 45%. Based on our findings, we design a prototypical registration and screening system with integrated information flow that can be used for dispensing, monitoring and assessing mass vaccination. The system connects to the local Immunization Information System and electronic medical record systems. The design is flexible and adaptable for different types of medical countermeasures, and is suitable for regional public health departments. Our approach bridges research and public health informatics in a practical way, demonstrating how data can guide both system design and public health response planning.
Researchers have developed pharmacogenomics datasets for various purposes, such as biomarker identification, yet drug response prediction models often underperform due to dataset inconsistencies. These variations arise from inter-tumoral heterogeneity, experimental conditions, and cell subtype complexity, limiting model generalizability. To address this, we propose a computational model based on Aggregated Learning (AL) to enhance drug response prediction by learning from inconsistencies across multiple datasets. Our model minimizes discrepancies by training on overlapping inconsistent data points from three pharmacogenomic datasets-CCLE, GDSC2, and gCSI. Compared to four baseline methods-Selecting Better (SB), Result Average (RA), Combining Data (CD), and Model Average (MA)-our approach achieved superior performance with lower Mean Absolute Error (MAE) scores: 0.090 (CCLE-GDSC), 0.096 (CCLE-gCSI), and 0.081 (GDSC-gCSI). These results demonstrate that addressing inconsistencies enhances prediction accuracy and generalizability, making our model a promising solution for robust drug response predictions.
This work introduces the Sequential Multiple Instance Learning (SMIL) framework, addressing the challenge of interpreting sequential, variable-length sequences of medical images with a single diagnostic label. Diverging from traditional MIL approaches that treat image sequences as unordered sets, SMIL systematically integrates the sequential nature of clinical imaging. We develop a bidirectional Transformer architecture, BiSMIL, that optimizes for both early and final prediction accuracies through a novel training procedure to balance diagnostic accuracy with operational efficiency. We evaluated BiSMIL on three medical image datasets to demonstrate that it simultaneously achieves state-of-the-art final accuracy and superior performance in early prediction accuracy, requiring 30-50% fewer images for a similar level of performance compared to existing models. Additionally, we introduce SMILU, an interpretable uncertainty metric that outperforms traditional metrics in identifying challenging instances.
Intimate Partner Violence (IPV) remains a significant global health issue with severe consequences ranging from physical injury to death, with rates rising in recent years. Prediction of recidivism is critical for prevention and treatment. Using data from a four-year clinical study, we develop interpretable machine-learning models to identify features for physical assault recidivism among IPV offenders. To standardize clinician-assigned severity scores and address non-linear associations, we apply filtered target encoding, which reduces subjectivity and bias in assessment. We find that combining self-reported and partner-reported variables enhances predictive power. Through feature importance analyses, we identify factors associated with lower recidivism risk, including decreased substance use and avoiding partner contact, while separation processes correlate with higher reoffending likelihood. These findings advance IPV risk assessment by providing a deeper understanding of risk factors critical for improving treatment effectiveness and addressing disparities in IPV management.
Adolescents and young adults with kidney transplants face unique challenges as they transition toward independent self-management. These youth need tools that not only support skill-building but also foster reflection, a key component of self-management. To address this need, we created a digital prototype designed to help users reflect on their transplant journey through storytelling and evaluated it in a study with 23 participants (13 youth and 10 caregivers) using. Participants found value in engaging with the prototype, which helped them reflect on their past experiences, gain insight into their current journey, and envision their future. Youth, in particular, reported increased self-awareness and confidence in managing their health. Based on these findings, we present design recommendations for future digital health tools aimed at supporting self-management in youth with chronic conditions.
With data considered as the 'oxygen' of public health, the Data Modernization Initiative (DMI) to enhance the public health data and information infrastructure is critical. The DMI Stories from the Field features data modernization from public health agencies to highlight success/progress. These stories (n=241) were analyzed, with outbreak response, information systems capacity, epidemiology/laboratory capacity being some of the common topics. A total of 199 codes across DMI stories were organized into 7 themes and the top 3 codes were communication, collaboration and public health agencies. Key takeaways and next steps were identified and validated with expert input across people, product, process and partnership categories and people factor was critical along with funding/sustainability. Ongoing DMI stories and future studies for evaluating impact are recommended. DMI stories are a great option to communicate the projects and impact of DMI to a larger public audience and garner support for this vital endeavor.
In the last decade, varied state-level policies on contraception access have highlighted the importance of large-scale public health datasets in assessing the impact of these policies on reproductive healthcare access. This study uses PRAMS Phase 8 (2016-2022) data to examine predictive factors of postpartum birth control use, hypothesizing that state policies impact contraception uptake and barriers, particularly regarding the expansion of immediate postpartum long-acting reversible contraception (LARC) reimbursement policies. Two distinct logistic regression models were constructed, and the inclusion of state as a covariate significantly reduced residual deviance (ΔDeviance = 13.696, p=0.0002). This finding indicates that state of residence is a statistically significant predictor of postpartum birth control usage. This study underscores the significant impact of state-level and institutional policies on birth control usage and LARC uptake, emphasizing the need for informed policy changes and patient-centered strategies to address disparities and improve postpartum reproductive health outcomes.
Introduction: Despite extensive literature on alert fatigue, gaps remain in understanding its impact. We examined drug alert volume and override rates across provider roles to inform future research. Methods: We retrospectively analyzed drug allergy alerts (DAA) and drug-drug interaction (DDI) alerts in 2023 at a large academic medical center. Alert volume and override rates were compared across providers with prescribing authority. Results: Among 1,799 providers, 196,225 alerts were generated with an average of 0.42 alerts per clinician day. Advanced practice providers (APPs) received significantly more alerts per day than residents or attending physicians. Most providers (88%) saw fewer than one alert daily. Override rates increased with higher alert burden (98.6% for >5 alerts/day vs. 94.5% for 1-5 and 92.6% for <1; p<0.001). Conclusion: Alert fatigue may not be observed in all cases when analyzed at the provider level. Future research should explore other alert characteristics, measurements, and provider's perceptions.
The emergence of novel infectious pathogens challenges early-phase modeling of disease transmission due to limited, low-quality data and an incomplete understanding of the pathogen. Additionally, regional variations in outbreaks necessitate models that incorporate local dynamics. We present an early-phase local model that leverages constrained public health data, primarily infection counts and aggregated regional characteristics, to study disease transmission dynamics. To address data limitations and potential model misspecifications, we incorporate a quasi-likelihood approach with a flexible error term. Furthermore, we introduce an online estimator that enables real-time data updates, supported by an iterative algorithm for parameter estimation. We applied this method to early COVID-19 data, analyzing infection counts and county-level risk factors from more than 800 U.S. counties to predict disease spread and assess the impact of social behavior, demographics, and vaccination coverage on disease transmission. This framework improves early outbreak analysis and informs local pandemic response under suboptimal data conditions.
Cardiovascular event adjudication is essential in clinical trials but relies on manual chart review that is slow, variable, and expensive. We present a two-stage framework that automates adjudication of cardiovascular deaths using large language models (LLMs). First, a few-shot LLM extracts structured evidence (event, span, negation, date) from unstructured clinical documents. Second, a Tree-of-Thoughts adjudicator aligns its reasoning with clinical endpoint committee (CEC) guidelines to classify deaths as cardiovascular or non-cardiovascular and produce an auditable rationale. On Lilly clinical-trial data, extraction achieved precision 0.96, recall 0.71 (F1 0.82), and adjudication attained 0.68 accuracy (GPT-4 ToT), outperforming a summarizer-plus-adjudicator baseline. We introduce CLEART, a rubric-based automated score that quantifies rationale quality across clarity, consistency, detail, guideline adherence, relevance, and timeline accuracy (overall 0.67), highlighting temporal reasoning and relevance as key areas for improvement. This approach can reduce adjudication time and variability while increasing transparency.
This modified explanatory sequential mixed methods study sought to inform redesign of nursing notes in the electronic health record. In the context of OpenNotes and patient and family access to nursing notes via the inpatient portal, redesigning nursing notes offers an opportunity to enhance family-centered care delivery and reduce nurses' documentation burden. We analyzed data on note views via the inpatient portal for 258,841 nursing notes; annotated the contents of 100 nursing notes; and conducted interviews with 18 families and 8 nurses. Our findings support recommendations for more specific care plans, eliminating redundancies, and emphasizing nursing care and expertise otherwise absent from the patient chart. The results of this descriptive study lay the groundwork for pilot testing new nursing note structures.
Artificial intelligence and machine learning are transforming healthcare by improving clinical risk predictions and diagnostic precision. However, their performance can be compromised by data drifts due to changes in patient populations and evolving clinical practices. This study investigated performance drift in models predicting Acute Kidney Injury (AKI) using electronic health records from 249,749 inpatient encounters over ten years, analyzing performance across both the overall population and nine subgroups with unique health profiles. To mitigate the performance drift, we implemented two model updating strategies: an Overall Population Update (OPU) and a Specific Subgroup Update (SSU). Our results demonstrated significant reductions in drift, with OPU increasing the average area-under-the-precision-recall-curve (AUPRC) by 0.14 in the overall population and 0.11 across subgroups, and SSU improving the average AUPRC by 0.10 among subgroups. These findings highlight the importance of continuous model surveillance and adaptive updates to maintain reliable predictive performance in dynamic clinical environments.
Acute kidney injury (AKI) is a severe condition in the ICU, where early prediction is crucial for timely intervention and prevention. Traditional machine learning (ML) models lack interpretability, which limits real-world applicability. We propose AKI-Detector, a novel multi-agent framework that integrates structured electronic health records (EHR)-based ML models, large language models (LLMs), and retrieval-augmented generation (RAG) to enhance clinical reasoning, accuracy, and interpretability of AKI prediction. The proposed AKI-Detector mitigates LLM hallucinations by integrating ML models and bridges the gap between algorithmic output and clinically interpretable reports. Evaluated on ICU data from MIMIC-IV, AKI-Detector outperformed ML models such as CatBoost and GRU, and achieved an accuracy of 0.827, precision of 0.672, recall of 0.542, and F1-score of 0.600 on the test cohort, demonstrating balanced and reliable predictive performance. This work highlights the promise of real-world big data and LLM-powered multi-agent systems to support trustworthy and explainable AI for clinical prediction.
Ventilator-Associated Pneumonia (VAP) significantly impacts critical care outcomes, yet prediction models often over-look healthcare data's meronomic structure. Using MIMIC-III data, we developed a multi-source extraction approach integrating structured data with clinical notes, identifying 679 VAP and 3,207 non-VAP cases. We compared four data splitting strategies: Ventilator Session-Based Split, Ventilator Session-Based Split on Single ICU Stays, Hospital Admission-Based Split, and Hospital Admission-Based Split on Single ICU Stays. Evaluating four machine learning models revealed that conventional random splitting yielded moderately high performance (AUROC: 76-81%) while restricting to single ICU stays surprisingly improved performance (AUROC: 86-87%). Admission-based approaches showed realistic results (AUROC: 72-76%). Feature analysis identified mechanical ventilation hours, systolic blood pressure, and urine counts as consistently important predictors. These findings demonstrate that robust VAP prediction requires evaluation frameworks respecting healthcare data's meronomic nature.
A challenge in utilizing electronic health record data for artificial intelligence models is contextualization, including understanding differences between missing data and missed care. Our team aims to develop knowledge graphs and computational models that account for these contexts, such as when data is missing (nurses being unable to document), but acceptable nursing care was delivered. We developed evaluation guidelines for intravenous and subcutaneous insulin management to establish a binary variable derived from EHR data representing minimally acceptable safe and quality nursing care for use in computational modeling. These guidelines were developed by our nurse informatics team based on best practices and validated by three nurse subject matter experts. The resulting evaluation guidelines are agnostic to institutional policies and focus on evaluating minimally acceptable safe and quality levels of care to inform inferences about missing data versus missed nursing care. Future work includes data-driven validations and expanding to other clinical scenarios.
Electronic cigarette (vaping) usage in the U.S. has steadily increased, raising significant public health concerns. Extensive research demonstrates various negative health outcomes associated with vaping. However, many potential harms remain understudied, especially those directly reported by users. Social media platforms such as Reddit offer rich, real-time sources of unfiltered personal accounts, presenting a unique opportunity to explore health outcomes beyond traditional clinical research. In this study, we systematically investigated potential negative health outcomes (NHOs) by analyzing millions of posts and comments from 15 active vaping-related subreddits in 2019. Employing robust data-driven methodologies, including advanced natural language processing (NLP) techniques such as sentiment analysis, UMLS tagging, and topic modeling, we identified distinct patterns of vaping-related health concerns. Our findings highlight the value of user-generated content for early detection of emerging risks, guiding clinicians, policymakers, and public health initiatives aimed at mitigating vaping-related harms, particularly among younger populations.
American Indian and Alaska Native (AI/AN) communities not only face significant health disparities but are often underrepresented in health research dissemination. Existing communication tools may fail to effectively reach these communities in culturally relevant and accessible ways, limiting their ability to benefit from critical health research. We co-designed and evaluated a prototype for health research results dissemination for AI/AN communities. We created and evaluated the prototype drawing from previous co-design workshops with AI/AN people. 38 participants completed an evaluation providing feedback for further iteration and highlighting key features such as search functionality, ease of use, visual and interactive elements, and content accessibility. Participants emphasized the importance of community connection, educational resources, and personalized experiences. We feature an alternative design approach we call Indigenous Community-Centered Design to create more accessible and engaging health research communication tools for AI/AN communities, fostering stronger connections and more accurate research representation.