Neurodevelopmental disorders (NDD) are a wide and heterogenous group of conditions due to impaired brain development, orchestrated by the crosstalk between genome and environment. Dynamic chromatin regulation during cortical development is fundamental, and chromatin remodelers are critical determinants of this process. Recently, numerous chromatin remodeling genes have been implicated in NDDs. By altering genes’ epigenetic state, mutated chromatin remodelers disrupt the spatiotemporal regulation of gene expression during development, potentially leading to severe consequences. The Remodeling and Spacing Factor 1 (RSF1) gene encodes a ubiquitous nuclear protein involved in chromatin remodeling, crucial for processes such as DNA transcription, replication, and repair. In this study, we identified by gene matching (n = 7) and literature search (n = 4) eleven unrelated individuals harboring de novo or inherited from a symptomatic parent heterozygous variants in RSF1. All individuals had an NDD, whether intellectual disability, autism spectrum disorder or developmental delay. From the seven individuals with detailed clinical information, unspecific and inconsistent associated features were described, including cranio-facial morphological features, musculoskeletal, digestive, vision, tone, epilepsy and brain MRI anomalies. Our data support the hypothesis that RSF1 is important for brain development and a novel candidate gene for syndromic NDDs.
Artificial intelligence (AI) has been used in many areas of medicine, and large language models (LLMs) have shown potential utility for various clinical applications. However, to determine if LLMs can accelerate the pace of genetic diagnosis and discovery, we examined whether recently developed LLMs (Med-PaLM 2 and Gemini) could assist in solving four types of genetic problems with sequentially increasing complexity. First, in response to free-text input, Med-PaLM 2 correctly identified murine genes with experimentally verified causative genetic factors for six previously studied murine models of biomedical traits. Second, Med-PaLM 2 identified a novel causative murine genetic factor for spontaneous hearing loss that was validated using knock-in mice. Third, we developed a retrieval and grounding pipeline that enabled Gemini 2.5 Pro to analyze large lists of genes, which contained genetic variants that were identified in the genomic sequences of 20 human subjects with hearing loss, and demonstrated that it can assist in identifying causative genetic factors for hearing loss. Fourth, we modified the genetic analysis pipeline to enable Gemini 2.5 Pro without any task-specific fine-tuning to identify causative genetic factors for six subjects with rare genetic diseases, which required 14 to 34 different terms to describe their multi-faceted symptom complexes. These results demonstrate that an AI pipeline can facilitate genetic diagnosis and discovery in mice and humans.
Purpose:While heterozygous de novo missense variants in the microtubule-binding GAR domain of Microtubule-actin cross-linking factor 1 (MACF1) cause Lissencephaly 9 with Complex Brainstem Malformations [MIM #618325], the phenotypic impact of variants outside this domain remains unclear. Methods:Through collaborative efforts, we assembled a cohort of 10 affected individuals from 8 unrelated families with either biallelic or monoallelic non-GAR domain MACF1 variants who exhibit partially overlapping yet unique phenotypic traits. Combined with previously reported cases, we analyzed genotype and phenotype data from 29 individuals using Human Phenotype Ontology (HPO)-based unsupervised hierarchical clustering. Results:Clustering revealed two distinct phenotypic signatures, suggesting domain-specific effects. Variants outside the GAR domain associate with broader neurodevelopmental phenotypes and variable craniofacial and skeletal expressivity. Additionally, enrichment analysis (p < 0.001) using OMIM HPO sets supported these findings. In contrast to the GAR domain's strong correlation with lissencephaly and brainstem malformations, biallelic non-GAR domain MACF1 variants were linked to diverse developmental anomalies. Conclusion:These results expand the phenotypic spectrum of MACF1-related disorders and highlight the relevance of domain-specific variant effects. Comprehensive genetic and phenotypic assessments are essential for understanding the role of MACF1 in development, informing diagnosis, and guiding future research on cytoskeletal regulation in neurodevelopment.
Background Well-recognized challenges in rare disease diagnosis include limited awareness of rare diseases among healthcare providers and barriers to accessing genetic testing. Less well understood are the ways in which communication between parents of undiagnosed children and providers may impact access to diagnosis, as well as quality of care broadly. We sought to characterize key dynamics of communication between parents of undiagnosed children and healthcare providers during the diagnostic odyssey. Methods Parents of undiagnosed children undergoing genomic sequencing were recruited from clinical and research settings and Facebook groups. Participants completed up to three sequential, in-depth interviews. Data were analyzed inductively to identify key themes. Results Parents (n = 36) identified three key dimensions of their experiences communicating with providers during the diagnostic odyssey, including examples of both effective and challenging communication related to: 1) providers' availability and responsiveness; 2) trust and validation of their concerns by providers; and 3) communication across multiple providers. Parents also described employing divergent strategies, such as increased persistence and advocacy, or minimized communication and resignation, in response to challenges. Conclusions Our study identified ways in which parent-provider communication can facilitate or hinder access to diagnosis and care for children with undiagnosed diseases. However, communication challenges were not universal, suggesting opportunities for intervention. Additional research is needed to identify interventions to improve parent-provider interactions during the diagnostic odyssey and to systematically evaluate the impact on time to diagnosis, access to care and patient health outcomes.
Rare structural variants (SVs)-insertions, deletions, and complex rearrangements-can cause Mendelian disease, yet they remain difficult to accurately detect and interpret. We sequenced and analyzed Oxford Nanopore Technologies long-read genomes of 68 individuals from the undiagnosed disease network (UDN) with no previously identified diagnostic mutations from short-read sequencing. Using our optimized SV detection pipelines and 571 control long-read genomes, we detected 716 long-read rare (MAF < 0.01) SV alleles per genome on average, achieving a 2.4× increase from short reads. To characterize the functional effects of rare SVs, we assessed their relationship with gene expression from blood or fibroblasts from the same individuals and found that rare SVs overlapping enhancers were enriched (LOR = 0.46) near expression outliers. We also evaluated tandem repeat expansions (TREs) and found 14 rare TREs per genome; notably, these TREs were also enriched near overexpression outliers. To prioritize candidate functional SVs, we developed Watershed-SV, a probabilistic model that integrates expression data with SV-specific genomic annotations, which significantly outperforms baseline models that do not incorporate expression data. Watershed-SV identified a median of eight high-confidence functional SVs per UDN genome. Notably, this included compound heterozygous deletions in FAM177A1 shared by two siblings, which were likely causal for a rare neurodevelopmental disorder. Our observations demonstrate the promise of integrating long-read sequencing with gene expression toward improving the prioritization of functional SVs and TREs in rare disease patients.
RNA sequencing has improved the diagnostic yield of individuals with rare diseases. Current analyses predominantly focus on identifying outliers in single genes that can be attributed to cis-acting variants within the gene locus. This approach overlooks causal variants with trans-acting effects on splicing transcriptome wide, such as variants impacting spliceosome function. We present a transcriptomics-first method to diagnose individuals with rare diseases by examining transcriptome-wide patterns of splicing outliers. Using splicing outlier detection methods (FRASER and FRASER2), we characterized splicing outliers from whole blood for 385 individuals from the Genomics Research to Elucidate the Genetics of Rare Diseases (GREGoR) and Undiagnosed Diseases Network (UDN) consortia. We examined all individuals for excess intron retention outliers in minor intron-containing genes (MIGs). Minor introns, which account for 0.5% of all introns in the human genome, are removed by small nuclear RNAs (snRNAs) in the minor spliceosome. This approach identified five individuals with excess intron retention outliers in MIGs, all of whom were found to harbor rare, bi-allelic variants in minor spliceosome snRNAs. Four individuals had rare, compound heterozygous variants in RNU4ATAC, which aided the reclassification of four variants. Additionally, one individual had rare, highly conserved, compound heterozygous variants in RNU6ATAC that may disrupt the formation of the catalytic spliceosome, suggesting it is a gene associated with Mendelian disease. These results demonstrate that examining RNA-sequencing data for transcriptome-wide signatures can increase the diagnostic yield of individuals with rare diseases, provide variant-to-function interpretation of spliceopathies, and uncover gene-disease associations.
Rare genetic diseases constitute a diverse group of disorders that affect 3.5% to 5.9% of people globally. Whole-genome sequencing (WGS) is a powerful tool that can potentially identify the genetic etiologies of these diseases. However, WGS does not always yield conclusive results, and may require periodic reanalysis with updated computational methods and database annotations to contribute to finding a diagnosis. This study aims to improve diagnostic yield of WGS by employing a targeted reanalysis approach with the Calypso Platform, a longitudinal genomic data management and genomic analysis platform.
Patients with undiagnosed and/or rare disorders frequently manifest dysmorphic and neurological features. There is a lack of information on the effectiveness of telehealth in the evaluation of these disorders. We thus compared an unassisted virtual physical examination (PE) with an in-person PE in undiagnosed individuals and also assessed participant telehealth satisfaction. Twenty-six individuals enrolled in the Undiagnosed Diseases Network study underwent an in-home synchronous virtual PE, and a subsequent in-person PE, by the same clinician. The participants completed surveys on telehealth usability and provider empathy. On PE, general appearance and craniofacial features showed near perfect agreement (κ = 0.81-1.00) between the telehealth and in-person evaluations. Specific components of the neurological examination demonstrated substantial agreement (speech, gait, coordination; κ > 0.61), whereas others had moderate agreement (muscle tone, strength; κ = 0.41-0.60), and a few had none to slight agreement (skin; κ < 0.20). Some systems could not be examined in the virtual PE. Importantly, features relevant to diagnosis or management were missed on the virtual PE in only 2/26 individuals. The participants were satisfied with the quality of the telehealth interaction, as well as empathy demonstrated by providers in the virtual interface. Telehealth is effective for PEs in undiagnosed diseases and is acceptable to affected individuals.
Long read sequencing offers benefits for the detection of structural variation in Mendelian disease. Here, we applied a new technology that generates contiguous long reads via tagmentation and sequencing by synthesis to a small cohort of patients with undiagnosed disease from the Undiagnosed Diseases Network. We first compare sequencing from the HG002 benchmark sample from Genome In A Bottle using nanopore sequencing (R10.4.1, duplex reads, Oxford Nanopore), single molecule real time sequencing (Revio SMRT cell, Pacific Biosciences) and complete long read sequencing (S4 flowcell, Novaseq, Illumina). Coverage was 33-35x across platforms. Read length N50 was 6.5kb (ICLR), 16.9kb (SMRT), and 33.8kb (ONT). We noted small differences in single nucleotide variant F1 scores across long read technologies with single nucleotide variant F1 scores (0.985-0.999) exceeding indel scores (0.78-0.99) and structural variant scores (0.74-0.96). We applied CLR sequencing to seven undiagnosed patients. In one patient, we detected and prioritized a novel 16kb intragenic duplication encompassing exons 5 and 6 in EHMT1. Resolution of the breakpoints and examination of flanking sequences revealed that the duplication was present in tandem and was predicted to result in a frameshift of the amino acid sequence and an early termination codon. It resulted in a diagnosis of Kleefstra syndrome. The variant was confirmed with targeted EHMT1 clinical testing and detected via nanopore and SMRT sequencing. In summary, we report the early clinical application of complete long read sequencing to a small cohort of undiagnosed patients. ### Competing Interest Statement Euan A. Ashley has received support in kind from Illumina, PacBio, and Nanopore. ### Funding Statement This study was funded by the Undiagnosed Diseases Network. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Participants sequenced in this study were enrolled in the Undiagnosed Diseases Network (UDN) at Stanford Medicine and provided informed consent. The study was granted ethical approved by the central IRB at the National Institutes of Health. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes The data on HG002 Genome in a Bottle sample is available publicly and referenced in the Methods section. The Ilumina complete long read sequencing data will be available in dbGaP in accordance with Undiagnosed Diseases Network data sharing policies. As data is deposited from the Undiagnosed Disease Network Data Management Coordinating Center Gateway Database to dbGaP, the complete long read sequencing data is not immediately available. Aggregate short read genome sequencing data on all UDN participants can be accessed via dbGaP Study Accession: phs001232.v5.p2.
Transcriptomics is a powerful tool for unraveling the molecular effects of genetic variants and disease diagnosis. Prior studies have demonstrated that choice of genome build impacts variant interpretation and diagnostic yield for genomic analyses. To identify the extent genome build also impacts transcriptomics analyses, we studied the effect of the hg19, hg38, and CHM13 genome builds on expression quantification and outlier detection in 386 rare disease and familial control samples from both the Undiagnosed Diseases Network and Genomics Research to Elucidate the Genetics of Rare Disease Consortium. Across six routinely collected biospecimens, 61% of quantified genes were not influenced by genome build. However, we identified 1,492 genes with build-dependent quantification, 3,377 genes with build-exclusive expression, and 9,077 genes with annotation-specific expression across six routinely collected biospecimens, including 566 clinically relevant and 512 known OMIM genes. Further, we demonstrate that between builds for a given gene, a larger difference in quantification is well correlated with a larger change in expression outlier calling. Combined, we provide a database of genes impacted by build choice and recommend that transcriptomics-guided analyses and diagnoses are cross referenced with these data for robustness.
Nucleic acid–sensing Toll-like receptors (TLR) 3, 7/8, and 9 are key innate immune sensors whose activities must be tightly regulated to prevent systemic autoimmune or autoinflammatory disease or virus-associated immunopathology. Here, we report a systematic scanning-alanine mutagenesis screen of all cytosolic and luminal residues of the TLR chaperone protein UNC93B1, which identified both negative and positive regulatory regions affecting TLR3, TLR7, and TLR9 responses. We subsequently identified two families harboring heterozygous coding mutations in UNC93B1, UNC93B1+/T93I and UNC93B1+/R336C, both in key negative regulatory regions identified in our screen. These patients presented with cutaneous tumid lupus and juvenile idiopathic arthritis plus neuroinflammatory disease, respectively. Disruption of UNC93B1-mediated regulation by these mutations led to enhanced TLR7/8 responses, and both variants resulted in systemic autoimmune or inflammatory disease when introduced into mice via genome editing. Altogether, our results implicate the UNC93B1-TLR7/8 axis in human monogenic autoimmune diseases and provide a functional resource to assess the impact of yet-to-be-reported UNC93B1 mutations.
RHOBTB2 encodes a member of the atypical Rho GTPase containing a GTPase domain and two tandem BTB domains. The BTB domains are involved in interacting with the Cullin3-dependent ubiquitin ligase complex, mediating ubiquitination, and recruiting substrates to the complex. Pathogenic de novo missense variants clustering in the BTB domains were reported to cause autosomal dominant developmental and epileptic encephalopathy 64 [DEE64; OMIM 618004]. DEE64 is a neurodevelopmental disorder characterized by onset of seizures within the first year of life, severe to profound intellectual disability, movement disorders, postnatal microcephaly, and nonspecific dysmorphic features.
PURPOSE:The function of FAM177A1 and its relationship to human disease is largely unknown. Recent studies have demonstrated FAM177A1 to be a critical immune-associated gene. One previous case study has linked FAM177A1 to a neurodevelopmental disorder in 4 siblings. METHODS:We identified 5 individuals from 3 unrelated families with biallelic variants in FAM177A1. The physiological function of FAM177A1 was studied in a zebrafish model organism and human cell lines with loss-of-function variants similar to the affected cohort. RESULTS:These individuals share a characteristic phenotype defined by macrocephaly, global developmental delay, intellectual disability, seizures, behavioral abnormalities, hypotonia, and gait disturbance. We show that FAM177A1 localizes to the Golgi complex in mammalian and zebrafish cells. Intersection of the RNA sequencing and metabolomic data sets from FAM177A1-deficient human fibroblasts and whole zebrafish larvae demonstrated dysregulation of pathways associated with apoptosis, inflammation, and negative regulation of cell proliferation. CONCLUSION:Our data shed light on the emerging function of FAM177A1 and defines FAM177A1-related neurodevelopmental disorder as a new clinical entity.
OBJECTIVE:ACTN2, encoding alpha-actinin-2, is essential for cardiac and skeletal muscle sarcomeric function. ACTN2 variants are a known cause of cardiomyopathy without skeletal muscle involvement. Recently, specific dominant monoallelic variants were reported as a rare cause of core myopathy of variable clinical onset, although the pathomechanism remains to be elucidated. The possibility of a recessively inherited ACTN2-myopathy has also been proposed in a single series. METHODS:We provide clinical, imaging, and histological characterization of a series of patients with a novel biallelic ACTN2 variant. RESULTS:We report seven patients from five families with a recurring biallelic variant in ACTN2: c.1516A>G (p.Arg506Gly), all manifesting with a consistent phenotype of asymmetric, progressive, proximal, and distal lower extremity predominant muscle weakness. None of the patients have cardiomyopathy or respiratory insufficiency. Notably, all patients report Palestinian ethnicity, suggesting a possible founder ACTN2 variant, which was confirmed through haplotype analysis in two families. Muscle biopsies reveal an underlying myopathic process with disruption of the intermyofibrillar architecture, Type I fiber predominance and atrophy. MRI of the lower extremities demonstrate a distinct pattern of asymmetric muscle involvement with selective involvement of the hamstrings and adductors in the thigh, and anterior tibial group and soleus in the lower leg. Using an in vitro splicing assay, we show that c.1516A>G ACTN2 does not impair normal splicing. INTERPRETATION:This series further establishes ACTN2 as a muscle disease gene, now also including variants with a recessive inheritance mode, and expands the clinical spectrum of actinopathies to adult-onset progressive muscle disease.
PURPOSE:Next-generation sequencing (NGS) has revolutionized the diagnostic process for rare/ultrarare conditions. However, diagnosis rates differ between analytical pipelines. In the National Institutes of Health-Undiagnosed Diseases Network (UDN) study, each individual's NGS data are concurrently analyzed by the UDN sequencing core laboratory and the clinical sites. We examined the outcomes of this practice.METHODS:A retrospective review was performed at 2 UDN clinical sites to compare the variants and diagnoses/candidate genes identified with the dual analyses of the NGS data.RESULTS:In total, 95 individuals had 100 diagnoses/candidate genes. There was 59% concordance between the UDN sequencing core laboratories and the clinical sites in identifying diagnoses/candidate genes. The core laboratory provided more diagnoses, whereas the clinical sites prioritized more research variants/candidate genes (P < .001). The clinical sites solely identified 15% of the diagnoses/candidate genes. The differences between the 2 pipelines were more often because of variant prioritization disparities than variant detection.CONCLUSION:The unique dual analysis of NGS data in the UDN synergistically enhances outcomes. The core laboratory provided a clinical analysis with more diagnoses and the clinical sites prioritized more research variants/candidate genes. Implementing such concurrent dual analyses in other genomic research studies and clinical settings can improve both variant detection and prioritization.