Diagnostic tools for rare diseases typically rely on curated gene-phenotype associations and static disease models, limiting their effectiveness in cases with atypical presentations or previously uncharacterized disorders. To address these limitations, we present SimPheny, a phenotype-first algorithm for gene prioritization that operates independently of documented gene-phenotype associations. SimPheny identifies phenotypically similar diagnosed patients by comparing an undiagnosed patient's disease presentation to a reference cohort of diagnosed cases, and returns gene hypotheses by matching the undiagnosed patient's candidate gene list to the causative genes of similar patients using a statistical scoring model. Evaluated in diagnosed probands from the Undiagnosed Diseases Network (UDN) with the true diagnostic gene blinded, SimPheny consistently ranked the diagnostic gene among the top five candidates, outperforming existing tools, particularly for genes with limited gene-phenotype association data. When applied to previously unsolved UDN cases, clinical review confirmed that SimPheny's high-confidence causative gene predictions were diagnostic in nearly half of the analyzed cases. As the size of the diagnosed reference cohort increases, SimPheny's diagnostic reach expands without sacrificing ranking performance. By leveraging real patient data rather than curated models, SimPheny provides a generalizable, scalable framework for improving diagnostic yield in rare disease cohorts.
Identifying critically ill newborns who will benefit from whole genome sequencing (WGS) is difficult and time-consuming due to complex eligibility criteria and evolving clinical features. The Mendelian Phenotype Search Engine (MPSE) automates the prioritization of neonatal intensive care unit (NICU) patients for WGS. Using clinical data from 2885 NICU patients, we evaluated the utility of different machine learning (ML) classifiers, clinical natural language processing (CNLP) tools, and types of Electronic Health Record (EHR) data to identify sick newborns with genetic diseases. Our results show that MPSE can identify children most likely to benefit from WGS within the first 48 h after NICU admission, a critical window for maximally impactful care. Moreover, MPSE provided stable, robust means to identify these children using many combinations of classifiers, CNLP tools, and input data types—meaning MPSE can be used by diverse health systems despite differences in EHR contents and IT support.
CONTEXT:DNA damage/repair gene variants are associated with both primary ovarian insufficiency (POI) and cancer risk. OBJECTIVE:We hypothesized that a subset of women with POI and family members would have increased risk for cancer. DESIGN:Case-control population-based study using records from 1995 to 2022. SETTING:Two major Utah academic health care systems serving 85% of the state. SUBJECTS:Women with POI (n = 613) were identified using International Classification of Diseases codes and reviewed for accuracy. Relatives were linked using the Utah Population Database. INTERVENTION:Cancer diagnoses were identified using the Utah Cancer Registry. MAIN OUTCOME MEASURES:The relative risk of cancer in women with POI and relatives was estimated by comparison to population rates. Whole genome sequencing was performed on a subset of women. RESULTS:Breast cancer was increased in women with POI (OR, 2.20; 95% CI, 1.30-3.47; P = .0023) and there was a nominally significant increase in ovarian cancer. Probands with POI were 36.5 ± 4.3 years and 59.5 ± 12.7 years when diagnosed with POI and cancer, respectively. Causal and candidate gene variants for cancer and POI were identified. Among second-degree relatives of these women, there was an increased risk of breast (OR, 1.28; 95% CI, 1.08-1.52; P = .0078) and colon cancer (OR, 1.50; 95% CI, 1.14-1.94; P = .0036). Prostate cancer was increased in first- (OR, 1.64; 95% CI, 1.18-2.23; P = .0026), second- (OR, 1.54; 95% CI, 1.32-1.79; P < .001), and third-degree relatives (OR, 1.33; 95% CI, 1.20-1.48; P < .001). CONCLUSION:Data suggest common genetic risk for POI and reproductive cancers. Tools are needed to predict cancer risk in women with POI and potentially to counsel about risks of hormone replacement therapy.
Rapid genomic diagnostics in the Neonatal Intensive Care Unit represents a paradigm shift in medicine with increasing evidence of the utility of early diagnosis, impacting management. The goal of the Utah NeoSeq Project was to implement and evaluate a multidisciplinary and longitudinal rapid sequencing program while transitioning to CLIA-certified sequencing. Enrollment of 65 infants resulted in 26 (40%) with a diagnostic variant(s) and 7 (11%) harboring a strong candidate. This includes re-analyses resulting in four additional diagnoses. Parental surveys indicated that 7% (4/59) of parents had a decisional conflict after consent, and 3% (2/59) experienced decisional regret after the results. Fifty-two provider surveys were conducted. Seventy-nine percent (41/52) of results and 86% (19/22) of diagnostic results were “very useful” or “useful” and associated with management changes. The NeoSeq Project demonstrates that a multidisciplinary collaborative approach to diagnosis is feasible. We have developed a generalizable, collaborative protocol that addresses the need for expedited genetic evaluation with emerging technologies.
Sudden Unexpected Infant Death (SUID), the third leading cause of infant death, has increasing incidence and multifactorial etiology. Identification of preventative interventions has hitherto been hindered by etiologic studies limited to genetic or environmental effects in isolation. Here we report a multifactorial genome x environment analysis of SUID risk. Births in San Diego County California from 2005-2018 were linked to hospital discharge summaries and death files, yielding 212 SUID cases and 620,392 infants alive at age 1 year. Whole genome sequencing (WGS) identified probable and possible genetic etiologies in 16% and 48% of SUID cases, respectively. Genetic risks were extremely heterogeneous with 144 loci contributing 173 risks in 57% of SUID cases. Genetic risk was very strong (Prevalence Risk Ratio, PRR >99) or strong (PRR 3.7 - 99) in 12% and 34% of SUID cases, respectively. Six of sixteen significant environmental risks lost significance when SUID cases without strong or very strong genetic risk were compared with infants alive at age 1 year, while SUID risk associated with prenatal cannabis increased from adjusted hazard ratio (aHR) 3.7 to 6.0, other substance abuse from aHR 2.6 to 3.5, and black race from aHR 1.9 to 2.5. Thus, genome x environment analysis of a large cohort unveiled etiologic heterogeneity and hidden SUID risks, highlighting cannabis and genetic diseases as strong risk factors. Since preventative or therapeutic interventions were available for 83% of genetic risks, newborn screening by WGS has potential for substantial SUID reduction. Educational campaigns for SUID should emphasize perinatal cannabis avoidance.
Background Identifying patients who would benefit from whole genome sequencing (WGS) is difficult and time-consuming due to complex eligibility criteria, lack of neonatologist familiarity with WGS ordering, and evolving clinical features. In previous work, we showed that MPSE, the Mendelian Phenotype Search Engine, can provide automated prioritization of probands for WGS while maintaining current diagnostic rates. MPSE is now in use in multiple hospital networks, but questions still surround how to best prioritize patients for WGS. Methods Here we use the clinical histories of 2,885 neonatal intensive care unit (NICU) admits from two institutions to explore further questions regarding how to best prioritize NICU admits for WGS. First, we ask if changes to the machine learning (ML) classifier and the clinical natural language processing (CNLP) tools used for generating patient phenotype descriptions might improve MPSE's performance. Second, we explore the utility of using alternative data types as inputs to MPSE. Lastly, we conduct a longitudinal analysis of MPSE's ability to identify probands for WGS. Results Eight different ML classifiers, five CNLP tools, and four previously untested alternative data types were used to train and validate MPSE models. MPSE achieved high predictive performance across multiple classifiers (max AUC=0.93), CNLP tools (max AUC=0.91), and input data types (max AUC=0.91). Longitudinal analysis of MPSE scores revealed a significant separation between cases/controls and diagnostic/non-diagnostic cases within 48 hours of NICU admission. Conclusions MPSE provides a highly flexible and portable framework for automated prioritization of critically ill newborns for WGS. We find that MPSE's performance is largely agnostic with respect to CNLP tools. Moreover, structured data such as ICD codes can serve as an effective alternative input to MPSE when access to clinical notes or CNLP pipelines is problematic. Finally, MPSE can identify children most likely to benefit from WGS within 48 hours of admission to the NICU, a critical window for maximally impactful care. ### Competing Interest Statement M.Y. is a co-founder and consultant for Fabric Genomics Inc.. M.R. is a shareholder of Fabric Genomics Inc.. E.F. is an employee of Fabric Genomics Inc.. B.M. and J.H. have received consulting fees and stock grants from Fabric Genomics Inc.. The remaining authors declare that they have no competing interests. ### Funding Statement The preparation of this manuscript was supported by a National Library of Medicine training grant (grant number T15LM007124), NIH grant UL1TR002550 from NCATS to E.J. Topol (with sub-award to Rady Children's Institute for Genomic Medicine), and the Warren Alpert Foundation. The Utah NeoSeq Project was funded by the Center for Genomic Medicine at the University of Utah Health, ARUP Laboratories, the Ben B. and Iris M. Margolis Foundation, the R. Harold Burton Foundation, and the Mark Miller Foundation. This work utilized resources and support from the Center for High Performance Computing at the University of Utah. The computational resources used were partially funded by the NIH Shared Instrumentation grant 1S10OD021644-01A1. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Library of Medicine or the National Institutes of Health. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The need for Institutional Review Board Approval at Rady Children's Hospital for the current study was waived as all data used from this project had previously been generated as part of IRB approved studies and none of the results reported in this manuscript can be used to identify individual patients. The studies from which cases were derived were previously approved by the Institutional Review Boards of Rady Children's Hospital. The University of Utah Institutional Review Board approved the use of human subjects for this research, under a waiver for the requirement to obtain informed consent. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study are available upon reasonable request to the authors. MPSE source code, pre-trained models, documentation, and synthetic datasets are available to the public on GitHub (https://github.com/Yandell-Lab/MPSE).
Genome-sequence-based newborn screening (gNBS) has substantial potential to improve outcomes in hundreds of severe childhood genetic disorders (SCGDs). However, a major impediment to gNBS is imprecision due to variants classified as pathogenic (P) or likely pathogenic (LP) that are not SCGD causal. gNBS with 53,855 P/LP variants, 342 genes, 412 SCGDs, and 1,603 therapies was positive in 74% of UK Biobank (UKB470K) adults, suggesting 97% false positives. We used the phenomenon of purifying hyperselection, which acts to decrease the frequency of SCGD causal diplotypes, to reduce false positives. Training of gene-disease-inheritance mode-diplotype tetrads in 618,290 control and affected subjects identified 293 variants or haplotypes and seven genes with variable inheritance contributing higher positive diplotype counts than consistent with purifying hyperselection and with little or no evidence of SCGD causality. With these changes, 2.0% of UKB470K adults were positive. In contrast, gNBS was positive in 7.2% of 3,118 critically ill children with suspected SCGDs and 7.9% of 705 infant deaths. When compared with rapid diagnostic genome sequencing (RDGS), gNBS had 99.1% recall. In eight true-positive children, gNBS was projected to decrease time to diagnosis by a median of 121 days and avoid life-threatening disease presentations in four children, organ damage in six children, ∼$1.25 million in healthcare cost, and ten (1.4%) infant deaths. Federated training predicated on purifying hyperselection provides a general framework to attain high precision in population screening. Federated training across many biobanks and clinical trials can provide a privacy-preserving mechanism for qualification of gNBS in diverse genetic ancestries.
Abstract Disclosure: E.B. Johnstone: None. B. Gorsi: None. E. Coelho: None. B. Moore: Consulting Fee; Self; Fabric Genomics. C. Chow: None. M. Yandell: Consulting Fee; Self; Fabric Genomics. Stock Owner; Self; Fabric Genomics. C.K. Welt: None. Context: A genetic etiology accounts for the majority of unexplained primary ovarian insufficiency (POI). Objective: We hypothesized a genetic cause of POI for a sister pair with primary amenorrhea.Design: The study was an observational study.Setting: Subjects were recruited at an academic institution.Subjects: Subjects were sisters with primary amenorrhea caused by POI, and their parents. Additional subjects included women with POI analyzed previously (n=291). Controls were recruited for health in old age or were from the 1000 Genomes Project (total n=233). Intervention: We performed whole exome sequencing (WES) and data were analyzed using the Pedigree Variant Annotation, Analysis and Search Tool (pVAAST), which identifies genes harboring pathogenic variants in families. We performed functional studies in a D. melanogaster model.Main Outcome: Genes with rare pathogenic variants were identified. Results: The sisters carried compound heterozygous variants in DIS3. The sisters did not carry additional rare variants that were absent in publicly available datasets. DIS3 knockdown in the ovary of D. melanogaster resulted in lack of oocyte production and complete infertility.Conclusions: Compound heterozygous variants in highly conserved amino acids in DIS3 and failure of oocyte production in a functional model suggest that mutations in DIS3 cause POI. DIS3 is a 3’ to 5’ exoribonuclease that is the catalytic subunit of the exosome involved in RNA degradation and metabolism in the nucleus. The findings provide further evidence that mutations in genes important for transcription and translation are associated with POI. Presentation Date: Saturday, June 17, 2023
Successful vaccines require adjuvants able to activate the innate immune system, eliciting antigen-specific immune responses and B-cell-mediated antibody production. However, unwanted secondary effects and the lack of effectiveness of traditional adjuvants has prompted investigation into novel adjuvants in recent years. Protein-coated microcrystals modified with calcium phosphate (CaP-PCMCs) in which vaccine antigens are co-immobilised within amino acid crystals represent one of these promising self-adjuvanting vaccine delivery systems. CaP-PCMCs has been shown to enhance antigen-specific IgG responses in mouse models; however, the exact mechanism of action of these microcrystals is currently unclear. Here, we set out to investigate this mechanism by studying the interaction between CaP-PCMCs and mammalian immune cells in an in vitro system. Incubation of cells with CaP-PCMCs induced rapid pyroptosis of peripheral blood mononuclear cells and monocyte-derived dendritic cells from cattle, sheep and humans, which was accompanied by the release of interleukin-1β and the activation of Caspase-1. We show that this pyroptotic event was cell–CaP-PCMCs contact dependent, and neither soluble calcium nor microcrystals without CaP (soluble PCMCs) induced pyroptosis. Our results corroborate CaP-PCMCs as a promising delivery system for vaccine antigens, showing great potential for subunit vaccines where the enhancement or find tuning of adaptive immunity is required.
CONTEXT:A genetic etiology accounts for the majority of unexplained primary ovarian insufficiency (POI).OBJECTIVE:We hypothesized a genetic cause of POI for a sister pair with primary amenorrhea.DESIGN:The study was an observational study. Subjects were recruited at an academic institution.SUBJECTS:Subjects were sisters with primary amenorrhea caused by POI and their parents. Additional subjects included women with POI analyzed previously (n = 291). Controls were recruited for health in old age or were from the 1000 Genomes Project (total n = 233).INTERVENTION:We performed whole exome sequencing, and data were analyzed using the Pedigree Variant Annotation, Analysis and Search Tool, which identifies genes harboring pathogenic variants in families. We performed functional studies in a Drosophila melanogaster model.MAIN OUTCOME:Genes with rare pathogenic variants were identified.RESULTS:The sisters carried compound heterozygous variants in DIS3. The sisters did not carry additional rare variants that were absent in publicly available datasets. DIS3 knockdown in the ovary of D. melanogaster resulted in lack of oocyte production and severe infertility.CONCLUSIONS:Compound heterozygous variants in highly conserved amino acids in DIS3 and failure of oocyte production in a functional model suggest that mutations in DIS3 cause POI. DIS3 is a 3' to 5' exoribonuclease that is the catalytic subunit of the exosome involved in RNA degradation and metabolism in the nucleus. The findings provide further evidence that mutations in genes important for transcription and translation are associated with POI.
BACKGROUND:Rapidly and efficiently identifying critically ill infants for whole genome sequencing (WGS) is a costly and challenging task currently performed by scarce, highly trained experts and is a major bottleneck for application of WGS in the NICU. There is a dire need for automated means to prioritize patients for WGS. METHODS:Institutional databases of electronic health records (EHRs) are logical starting points for identifying patients with undiagnosed Mendelian diseases. We have developed automated means to prioritize patients for rapid and whole genome sequencing (rWGS and WGS) directly from clinical notes. Our approach combines a clinical natural language processing (CNLP) workflow with a machine learning-based prioritization tool named Mendelian Phenotype Search Engine (MPSE). RESULTS:MPSE accurately and robustly identified NICU patients selected for WGS by clinical experts from Rady Children's Hospital in San Diego (AUC 0.86) and the University of Utah (AUC 0.85). In addition to effectively identifying patients for WGS, MPSE scores also strongly prioritize diagnostic cases over non-diagnostic cases, with projected diagnostic yields exceeding 50% throughout the first and second quartiles of score-ranked patients. CONCLUSIONS:Our results indicate that an automated pipeline for selecting acutely ill infants in neonatal intensive care units (NICU) for WGS can meet or exceed diagnostic yields obtained through current selection procedures, which require time-consuming manual review of clinical notes and histories by specialized personnel.
Adiponectin, encoded by ADIPOQ , is an insulin-sensitizing, anti-inflammatory, and renoprotective adipokine that activates receptors with intrinsic ceramidase activity. We identified a family harboring a 10-nucleotide deletion mutation in ADIPOQ that cosegregates with diabetes and end-stage renal disease. This mutation introduces a frameshift in exon 3, resulting in a premature termination codon that disrupts translation of adiponectin’s globular domain. Subjects with the mutation had dramatically reduced circulating adiponectin and increased long-chain ceramides levels. Functional studies suggest that the mutated protein acts as a dominant negative through its interaction with non-mutated adiponectin, decreasing circulating adiponectin levels, and correlating with metabolic disease.