PURPOSE:We describe a prospective cohort study (NCT05277116) conducted in phase IV of the electronic MEdical Records and GEnomics (eMERGE) Network to implement a multi-ancestry polygenic risk score for coronary heart disease (PRSCHD: PGS004696) and assess outcomes after return of results (RoR). METHODS:PRSCHD was considered alongside family history (FamHxCHD), monogenic risk from familial hypercholesterolemia (FH), and clinical risk factors, to return CHD risk as part of a Genome Informed Risk Assessment (GIRA) report. Participants with high PRSCHD (top 5th percentile) or FH received their results from study personnel, while participants with FamHxCHD were informed by mail/email. Results were placed in the electronic health record and communicated to the primary care provider. The primary outcome of initiation/intensification of lipid lowering therapy within 12 months after RoR is compared between participants with PRSCHD ≥95th percentile and those with PRSCHD 90th-94th percentile, using a regression discontinuity design. Secondary outcomes include ordering of screening tests, a new CHD diagnosis, and lifestyle changes. RESULTS:By April 2025, 20,421 adults were enrolled: mean age 50±15 years (range 18-75 years), 68% female, 50% belonging to health disparity groups, and 40% non-White by self-report. Prevalence of CHD, FamHxCHD, high PRSCHD and FH was 4.0%, 10.2%, 4.3% and 0.7%, respectively; 14.3% had at least one of the three CHD genetic risk factors and CHD risk estimates were highest in those who self-reported as Black. CONCLUSION:The prevalence of increased genetic risk for CHD was high and at least one of the three genetic risk factors for CHD was present in 14.2% of the cohort. Analyses are underway to assess outcomes after PRSCHD implementation in the context of FamHxCHD, FH, and clinical risk, across the age spectrum in a diverse cohort.
Polymorphisms thiopurine-S-methyltransferase (TPMT) and nudix hydrolase 15 (NUDT15) can increase the risk of azathioprine myelotoxicity, but little is known about other genetic factors that increase risk for azathioprine-associated side effects. PrediXcan is a gene-based association method that estimates the expression of individuals' genes and examines their correlation to specified phenotypes. As proof of concept for using PrediXcan as a tool to define the association between genetic factors and azathioprine side effects, we aimed to determine whether the genetically predicted expression of TPMT or NUDT15 was associated with leukopenia or other known side effects. In a retrospective cohort of 1364 new users of azathioprine with EHR-reported White race, we used PrediXcan to impute expression in liver tissue, tested its association with pre-specified phecodes representing known side effects (e.g., skin cancer), and completed chart review to confirm cases. Among confirmed cases, patients in the lowest tertile (i.e., lowest predicted) of TPMT expression had significantly higher odds of developing leukopenia (OR=3.30, 95%CI: 1.07-10.20, p=0.04) versus those in the highest tertile; no other side effects were significant. The results suggest that this methodology could be deployed on a larger scale to uncover associations between genetic factors and drug side effects for more personalized care.
Introduction: Recurrent psychiatric hospitalizations in youth are costly and burdensome to patients and families. While postdischarge planning rarely stratifies patients based on their risk of readmission, predictive modeling could identify at-risk youth to facilitate targeted intervention strategies. This study developed and validated machine learning models to predict psychiatric readmission in children and adolescents. Methods: This retrospective cohort study included 11,225 patients (20,339 psychiatric admissions) ≤ 18 years old with at least one psychiatric admission and an anxiety or depressive disorder diagnosis at any time during the study period from two tertiary-care academic medical centers. Machine learning models were developed to predict psychiatric readmission versus no readmission within 30, 90, and 180 days of discharge using electronic health record data from one institution. These models were internally validated at the development site and externally validated at the second institution. Results: Model development and internal validation using random forest models predicted 30-, 90-, and 180-day readmission with area under the receiver operating characteristic curves (AUROC) equal to 0.739, 0.742, and 0.746, respectively. The most important features identified across follow-up periods included previous psychiatric admissions, length of stay, age, and antipsychotic prescriptions. Models trained using a reduced set of 18 features obtained similar performance, with AUROCs ranging from 0.741 to 0.745. External validation of the models retrained using 12 features available at both institutions yielded AUROCs of 0.598, 0.634, and 0.657 for 30-, 90-, and 180-day readmission, respectively. Conclusion: The findings of this study suggest that machine learning models can identify children and adolescents at risk of psychiatric readmission.
Whether polygenic risk, monogenic familial hypercholesterolemia (FH), and family history (FamHx) are additively informative for coronary heart disease (CHD) risk prediction across self-identified race/ethnicity (SIRE) groups has not been established. In two diverse cohorts-Electronic Medical Records and Genomics (eMERGE) phase IV (eIV; n = 19,348) and All of Us (AoU; n = 239,645)-we quantified the associations of a polygenic risk score (PRSCHD), pathogenic/likely pathogenic variants in genes associated with FH, and FamHx with CHD and evaluated their incremental value when added to the pooled cohort equations (PCEs). CHD was defined as myocardial infarction, unstable angina, or coronary revascularization. We modeled associations with multivariable logistic regression (prevalent CHD in eIV) and Cox proportional hazards (incident CHD in AoU) and characterized predictive performance with the c-statistic and reclassification and decision-curve net benefits across actionable 10-year risk thresholds. The effects of PRSCHD and FamHx were independent and additive in both cohorts and consistent across White, Black, and Latino SIRE groups. In eIV, adding PRSCHD and FamHx to the PCE increased the c-statistic for prevalent CHD from 0.719 to 0.753 (p-diff = 9.1 × 10-3) and reclassified 18.8% of participants at the 7.5% 10-year threshold, yielding approximately 4 additional true-positive CHD identifications per 1,000 screened. Net benefit gains were observed between the 7.5% and 10% thresholds across all three SIRE groups. In conclusion, PRSCHD and FamHx were independently and additively associated with CHD across major SIRE groups in two diverse cohorts in the United States (US), motivating the addition of these factors to clinical risk algorithms.
Sharing biomedical research data can accelerate scientific discovery, leading funders and journals to increasingly mandate sharing. However, data openness must be balanced with protecting research participants from harm in an evolving legal and social landscape. Drawing on experiences from the Electronic Medical Records and Genomics (eMERGE-IV) Network—a US-based, multi-site consortium gathering genomic and medical data focused on underrepresented groups to refine disease risk prediction—we examine challenges in implementing data sharing that are “as open as possible, as closed as necessary.” Recent US legal developments, including the Dobbs decision and gender-affirming care bans, highlight the urgency of considering data-sharing risks and required the Network to rethink strategies to prevent individual- and group-level harms from genomic analyses. eMERGE-IV implemented several strategies to mitigate concerns, including cell suppression for race/ethnicity data and not extracting certain diagnostic codes from participants’ electronic health records. These decisions balanced immediate protection and long-term scientific benefits for relevant populations. Participant agreement to broad data sharing in informed consent is often required for research participation to make data as open as possible. No consent form, however, can define the terms of “as closed as necessary”—a construct that is subject to sociolegal changes across the life cycle of research studies. Providing protection requires robust data governance, including engagement with prospective and actual participants. The research enterprise must reconsider its consenting approach and develop transparent, inclusive governance structures responsive to evolving vulnerabilities while maintaining scientific progress. Public trust depends on the research enterprise successfully navigating these competing demands.
Rare disease diagnosis depends on expert reasoning that is scarce and difficult to transfer. Large language models rank the correct disease first in only 35.4
Incorporating genetic risk factors to assess health risk and inform screening is critical for advancing precision medicine. The Electronic Medical Records and Genomics (eMERGE) Network conducted a large-scale study returning genome-informed risk assessments (GIRAs) to 23,840 participants (ages 3-75) across ten clinical sites. Risk for 11 common conditions was assessed using polygenic risk scores (PRSs), monogenic variants, and family history, with results placed in the electronic health record and returned to participants. Non-high-risk and family history-only high-risk results were delivered via patient portal, secure email, or mail. High-risk results involving PRSs, monogenic variants, or BOADICEAs (integrated breast cancer risk scores) were attempted to be returned one-onone (1:1) via phone, video, or in person. Here, we evaluate the frequency and drivers of high-risk GIRAs, the feasibility of completing 1:1 returns, and factors associated with completing 1:1 delivery. Nearly 35% of participants (8,305) received high-risk results, most (76%) for a single condition. The most common triggers were family history and high PRS. Of those with high-risk GIRAs, 4,911 qualified for 1:1 return, with return completion rates equaling 78.5% for adults and 67.5% for children. The primary barrier to completing a 1:1 return session was the inability to contact participants. Among variables potentially impacting return success, homeownership, "good health," highest education level, and lack of health insurance all were significantly associated with successful 1:1 return with large effect sizes. This study demonstrates the feasibility of large-scale GIRA return in diverse clinical settings and highlights barriers that may impact equitable delivery of high-risk results.
Research applications of electronic health record (EHR) phenotypes require translating clinical definitions into executable EHR database queries, a labor-intensive process. We evaluated two frontier large language models across five phenotypes and three documentation modalities. Both models captured high-level logic from structured text but degraded markedly with diagram-only input. Error analysis revealed seven failure categories. Documentation, rather than model capability, was the primary bottleneck, reinforcing the need for standardization and expert oversight.
IntroductionDepression is a common psychiatric disorder and a leading cause of disability. Large-scale genomic studies have identified common variants associated with depression. However, researchers often rely on self-reported phenotypes, domain expertise-defined rules, and simple diagnostic codes (ICD9, ICD10) to identify depression participants, which suffers from inconsistent cohort definition and limited sample sizes. Thus, there is a lack of validated, efficient EHR phenotyping algorithms that precisely recognize depression cases.MethodsWe implemented a validated EHR phenotyping algorithm to construct a cohort of individuals with depression (11,532 cases and 39,631 controls, total n = 51,163) and conducted a genome-wide association study (GWAS) using this cohort. We validated the EHR-derived depression cohort using LDSC regression, comparing genetic similarities between our cohort and existing large meta-analyses. Top-ranked SNPs were selected and annotated to investigate downstream biological pathways and potential mechanisms that interfere with depression susceptibility.ResultsOur study reproduced previously identified genetic associations (PHF5A, KCNG2) with depression susceptibility. We also identified novel SNPs within the HLA region and IGVH region, suggesting an association between immune function and depression phenotype. We also demonstrated the robustness of the phenotyping algorithm through genetic correlation analysis (LDSC), showing a highly genetic similarity (rg = 0.8317, P = 1.7758e-11) between our cohort and large meta-analysis cohorts of major depressive disorder.ConclusionOur results demonstrate a robust validation of the EHR-based depression phenotyping algorithm using genetic analysis while providing novel genetic associations between depression and immune functions.
The Electronic Medical Records and Genomics (eMERGE) Network developed and implemented a genome-informed risk assessment (GIRA) to communicate genomic (polygenic risk scores [PRSs], integrated risk scores [IRSs], and monogenic results), clinical, and family history-based risk for 11 chronic diseases and provide recommended healthcare recommendations. GIRA reports have now been returned to 23,840 participants and their providers in a large prospective cohort study. We present here the study design and analysis framework for assessing the attributable impact of GIRA return. Pre-specified outcomes include (1) provider/participant adoption of recommended healthcare actions, (2) new diagnosis of disease, (3) treatment initiation/intensification, and (4) clinical outcomes (surrogate markers or clinical events). We assess outcomes in high risk vs. not-high-risk participants, adjusting for covariates. We evaluate the effect of PRS/IRS at pre-established high-risk thresholds using regression discontinuity (RD), a quasi-experimental method that mimics randomization near a cutoff, enabling estimation of causal effects and controlling for unobserved confounders. Monogenic and family history-based risk stratification are analyzed using logistic regression. With 23,840 participants and 12 months of follow-up, the study is powered to detect differences of 2%-11% with 80% power (α = 0.05 in the adoption outcome). Longer follow-up will be required to enable assessment of new disease diagnosis, treatment changes, and clinical outcomes. Through innovative RD analyses and defined outcomes and comparison groups, this study will provide new insights into the real-world clinical impact of genomic risk assessment, address critical evidence gaps, advance understanding of genomic medicine outcomes, and inform future research.
[This corrects the article DOI: 10.3389/fgene.2026.1818653.].
Urinary tract infections (UTIs) are traditionally viewed as environmentally driven, yet their inherited susceptibility remains largely unexplored. We conducted a cross-biobank genome-wide association study of recurrent UTIs in 1,860,836 individuals (213,869 cases and 1,646,967 controls). We identified 36 genetic susceptibility loci and performed tissue-based multi-omic mapping to prioritize candidate causal genes. UTI risk alleles preferentially modulated epithelial gene expression in kidney and bladder, converging on urinary epithelia structure and function. PSCA, encoding a secreted epithelial surface protein, emerged as the strongest candidate under genetic control; the gene product is constitutively secreted into the urine from kidney papilla and bladder epithelia, binds uropathogenic E. coli, and inhibits bacterial growth in vitro. Our findings define the polygenic architecture of UTIs and highlight the critical role of uroepithelial surface defenses, providing a new framework for host-directed, non-antibiotic interventions.
Current studies regarding the secondary use of electronic health records (EHR) predominantly rely on domain expertise and existing medical knowledge. A powerful representation approach can unleash the potential of discovering new medical patterns underlying the EHR. Here, we introduce an unsupervised method for embedding high-dimensional EHR data at the patient level to characterize heterogeneity in complex diseases and identify novel disease patterns linked to disparities in clinical outcomes. We applied this approach to 34,851 unique medical codes across 1,046,649 longitudinal patient events, including 102,740 patients in the Electronic Medical Records and GEnomics (eMERGE) Network. The model achieved strong predictive performance in predicting future disease (median AUROC = 0.87 within one year) and bulk phenotyping (median AUROC = 0.84). Notably, these patient embeddings revealed diverse comorbidity profiles and health outcomes, including distinct subtypes and progression patterns in colorectal cancer and systemic lupus erythematosus.
Objective:Phenotype libraries accelerate health research through the reuse of algorithms. However, multiple libraries exist with differing metadata standards. We describe the integration of libraries using shared conceptual standards to harmonize definitions. Materials and Methods:We identified a pilot set of definitions from the Observational Health Data Sciences and Informatics (OHDSI), Phenotype Knowledgebase (PheKB), and Health Data Research UK (HDR UK) phenotype libraries. Metadata were mapped to the Centralized Interactive Phenomics Resource (CIPHER) standard, integrated into the CIPHER library, and validated. Results:Twenty-seven phenotypes from OHDSI, 25 from PheKB, and 25 from HDR UK were mapped to the CIPHER standard and are available at https://phenomics.va.ornl.gov/. Discussion:We demonstrate the feasibility of integrating metadata across phenotype repositories and illustrate the CIPHER standard as a suitable initial framework for a global conceptual metadata model. Conclusion:Harmonizing phenotype metadata across libraries allows the reuse of definitions throughout the entire health data community.
Recent advances in medical vision-language models (VLMs) open up remarkable opportunities for clinical applications such as automated report generation, physician copilots, and uncertainty quantification. Despite their promise, medical VLMs raise serious security concerns. These include the risk of Protected Health Information (PHI) exposure, data leakage, and vulnerability to cyberthreats, concerns that are especially critical in hospital environments. Even when adopted for research or non-clinical purposes, healthcare organizations must exercise caution and implement safeguards. To address these challenges, we present MedFoundationHub, a graphical user interface (GUI) toolkit that: (1) enables physicians to manually select and use different models without programming expertise, (2) supports engineers in efficiently deploying medical VLMs in a plug-and-play fashion, with seamless integration of Hugging Face open-source models, and (3) ensures privacy-preserving inference through Docker-orchestrated, operating system agnostic deployment. MedFoundationHub requires only an offline local workstation equipped with a single NVIDIA A6000 GPU, making it both secure and accessible within the typical resources of academic research labs. To evaluate current capabilities, we engaged board-certified pathologists to deploy and assess five state-of-the-art VLMs (Google-MedGemma3-4B, Qwen2-VL-7B-Instruct, Qwen2.5-VL7B-Instruct, and LLaVA-1.5-7B/13B). Expert evaluation covered colon cases and renal cases, yielding 1,015 clinician-model scoring events. These assessments revealed recurring limitations, including off-target answers, vague reasoning, and inconsistent pathology terminology.
Rare diseases affect over 300 million people worldwide, yet patients often endure years-long diagnostic delays that limit timely intervention and trial opportunities. Computational rare disease recognition (RDR) remains constrained by knowledge resources that are often incomplete, heterogeneous, and dependent on extensive multi-disciplinary expert curation that cannot scale. Large language models (LLMs) applied directly for end-to-end diagnosis or disease discrimination face similar knowledge bottlenecks while also raising concerns around cost, reproducibility, and data governance. Here, we introduce GEN-KnowRD, a knowledge-layer-first framework that leverages LLMs to generate schema-guided rare disease profiles, systematically assesses their quality, and constructs a computable knowledge base (PheMAP-RD) for local deployment. GEN-KnowRD integrates this knowledge into lightweight inference pipelines for both general-purpose disease screening and specialized early discrimination from longitudinal electronic health records. Across six public benchmarks for general-purpose screen (9,290 patients spanning 798 rare diseases), GEN-KnowRD significantly improves disease ranking compared to a state-of-the-art, HPO-centered diagnostic framework (up to 345.8% improvement in top-1 success), advanced end-to-end LLM reasoning (up to 129.1% improvement), and a variant of GEN-KnowRD instantiated with expert-curated knowledge rather than LLM-generated profiles. In two real-world cohorts for early diagnosis of idiopathic pulmonary fibrosis (511 patients) as a use case, GEN-KnowRD also demonstrates robust discrimination performance gains, supporting effective RDR during the pre-diagnostic window. These findings demonstrate that repositioning LLMs from diagnostic reasoning to the knowledge layer-decoupling knowledge construction from patient-level inference-yields stronger RDR, while providing scalable, continuously updatable, and reusable infrastructure for diagnosis, screening, and clinical research across the rare disease landscape.
BACKGROUND:Survivors of hypoplastic left heart syndrome (HLHS), the most severe form of congenital heart disease, are at high risk for heart failure (HF). HF in early life is a major contributor to mortality in this vulnerable population. However, reliable approaches to identify infants at highest risk for early HF are currently lacking. OBJECTIVES:The purpose of this study was to evaluate whether ultra-rare variants in cardiomyopathy-associated genes are associated with HF risk in HLHS. METHODS:Neonates with HLHS were prospectively enrolled within the first 21 days of life at Duke University Health System. Children and adults with HLHS who were older than 21 days were enrolled from Duke University Health System and the University of North Carolina into an ambispective cohort. External HLHS cohorts from Nationwide Children's Hospital and Vanderbilt University Medical Center were evaluated to assess for reproducibility across institutions. Participants underwent genome sequencing, and ultra-rare variants in dilated cardiomyopathy-associated genes (minor allele frequency ≤0.01%) were evaluated. The primary outcome was HF, categorized as severe (ventricular assist device implantation, heart transplantation, or death) or medically managed (reduced systemic ventricular ejection fraction and/or HF diagnosis requiring initiation or escalation of HF therapy). Associations between variant status and HF risk were assessed using Cox regression. RESULTS:Among 35 neonates in the prospective cohort, 7 (20.0%) developed severe HF, 10 (28.6%) developed medically managed HF, and 18 (51.4%) remained HF free. The presence of a likely pathogenic/pathogenic variant was associated with a marked 9-fold increased risk of severe HF compared with genotype-negative individuals (P = 0.02). Most severe HF events occurred within the first month of life (67%). Similar associations between likely pathogenic/pathogenic variants and severe HF were observed in the Nationwide Children's Hospital and Vanderbilt University Medical Center cohorts (11- and 3-fold increased risk, respectively; all P < 0.05). Associations were attenuated in the ambispective cohort (all P > 0.05), which consisted of individuals significantly older than the prospective cohort (P < 0.0001). CONCLUSIONS:This study provides the first prospective evidence linking dilated cardiomyopathy-associated variants to early-onset HF in HLHS. These findings suggest that genetic variation may contribute to myocardial vulnerability in HLHS, highlighting the potential for genetic screening to enable early risk stratification and guide precision medicine approaches in this high-risk population.