The molecular pathways linking genetic variants to Parkinson's disease (PD) onset and progression remain incompletely defined; however, risk alleles in multiple genes, including GBA1, strongly implicate lipid metabolism. To systematically identify causal biomarker signatures, we analysed comprehensive metabolome profiles from blood plasma in 149 PD patients and 150 controls, along with complementary genetic, RNA-sequencing and metabolic data from other available clinical and pathologic cohorts. Using colocalization and summary-data-based Mendelian randomization, we tested whether expression and metabolic quantitative trait loci mediate the association between implicated genetic variants and PD risk. We further integrated differential metabolomics and proteomics from blood and brain to reveal pertinent mechanisms. We show that common PD risk variants at the serine palmitoyltransferase small subunit B (SPTSSB) locus, a key regulator of de novo sphingolipid biosynthesis, are associated with increased SPTSSB brain expression and elevated plasma ceramides. Additional analyses strongly support our hypothesis that a common SPTSSB causal variant is responsible for PD risk as well as the expression and metabolic quantitative trait loci. Multiple sphingolipids and fatty acid derivatives were perturbed in PD, and we identified both unique and shared features with the Alzheimer's disease metabolome. A PD acylcarnitine signature was further replicated in human post-mortem brain tissue, when comparing those with or without preclinical Lewy body pathology. Integrated analysis of complementary brain proteomic profiles revealed dysregulation of mitochondrial processes dependent on acylcarnitines, including fatty acid beta-oxidation, the tricarboxylic acid cycle and oxidative phosphorylation. Our results identify promising biomarkers and reveal a causal chain linking genetic variation to altered gene/protein expression, lipid dysmetabolism, and the manifestation of PD.
Nonsense-mediated decay (NMD) is a conserved RNA quality-control pathway that degrades transcripts containing premature termination codons. Because roughly a third of pathogenic variants in ClinVar can lead to truncated protein synthesis, predicting whether such transcripts undergo NMD is central to interpreting variant effects, yet the canonical 50-55 nucleotide rule explains only about half of observed outcome variability. Using paired whole-genome and RNA-sequencing from 10,306 individual samples in the Trans-Omics for Precision Medicine (TOPMed) program, we quantified NMD efficiency for 5,749 germline truncating variants via allele-specific expression and trained a gradient-boosting classifier, TrunCat, that distinguished NMD-sensitive from NMD-escape transcripts with ∼78% ROC-AUC (Receiver Operating Characteristic - Area Under the Curve). A reduced model using the ten features with the highest mean SHAP (SHapley Additive exPlanations) value as a measure of each feature's average contribution to predictions nearly matched this performance. Applied across large variant databases and a rare-disease cohort, the model produced NMD outcome predictions, with variants of uncertain significance showing higher predicted escape than pathogenic ones. This framework confirms the canonical rule, identifies non-canonical determinants, and offers a scalable resource for interpreting protein-truncating variants.
Background Parkinson's disease (PD) is a genetically complex disorder in which combinations of heterozygous risk variants may contribute to pathogenesis. Many PD risk loci encode lysosomal genes, such as GBA1, a common and potent risk factor, conferring at least a 5-fold increase. However, the mechanisms of GBA1 penetrance remain poorly understood. Methods Using Drosophila melanogaster, we performed a genetic interaction screen of lysosomal storage disorder (LSD) genes to identify dominant modifiers of Gba1b (fly homolog of GBA1). Age-dependent locomotor assessments, electroretinograms (ERG), transmission electron microscopy (TEM) analyses and quantification of dopaminergic (DA) neurons were used to assess the neurodegenerative phenotypes of double heterozygous animals. By combining immunostaining, lipidomics, metabolomics and pharmacological approaches we showed how partial loss of anne (fly homolog ofATP13A2) and Gba1b drives neurodegeneration. By interrogating genetic data from local and international PD cohorts we identified double heterozygous pathogenic variants in ATP13A2 and GBA1 in individuals with PD. Results We show that anne is expressed in neurons, whereas Gba1b is expressed in glia. Flies heterozygous for anne exhibit mild neurodegenerative phenotypes, and Gba1b strongly enhances this haploinsufficiency. Double heterozygous (Gba1b(T2A)/+;anne(T2A)/+) flies exhibit a slow and progressive neurodegeneration associated with accumulation and impaired acidification of lysosomes in photoreceptors and other neurons. Obvious morphological defects are first observed in glia at day 15 after eclosion and include vacuolization and neuronal detachment. These defects are accompanied by an elevation of glucosylceramide (GlcCer) and followed by loss of neuronal function and degenerative features by day 30. These phenotypes are neuronal activity-dependent. The neurodegenerative phenotypes are rescued by: ML-SA1, an agonist of the lysosomal TRPML1 channel that has been reported to promote lysosomal membrane trafficking; myriocin, a compound that inhibits GlcCer production; and DFMO, a drug which inhibits polyamine synthesis. Based on surveys of genetic data, we identify multiple PD cases harboring digenic variants in GBA1 and ATP13A2. Conclusions Our study reveals that partial loss of Gba1b in glia and anne in neurons synergistically disrupts lysosomal pH and neuron-glia GlcCer homeostasis, triggering neurodegeneration. Our results provide evidence that GBA1 penetrance is influenced by additional genetic modifiers, consistent with a putative digenic mechanism for GBA1-PD penetrance. These findings highlight lysosomal acidification, sphingolipid clearance, and polyamine regulation as critical intervention points in digenic PD.
WNT4 is a secreted protein that plays a critical role in the regulation of cell fate and embryogenesis. Biallelic variants in WNT4 have been linked to SERKAL syndrome, an autosomal recessive disorder characterized by 46,XX sex reversal and dysgenesis of the kidneys, adrenals, and lungs. SERKAL syndrome has only been described in a single consanguineous kindred with four affected fetuses. Additional features seen in a subset of affected fetuses included ventricular septal defect (VSD), congenital diaphragmatic hernia (CDH), and orofacial clefting (OFC). To determine if these additional features were likely to be caused by WNT4 deficiency, we used machine learning to compare WNT4 to genes known to cause VSD, CDH, and OFC. When compared to all RefSeq genes, WNT4's rank annotation scores for these congenital anomalies were 94%, 99%, and 98.5%, respectively, indicating a high level of similarity. We subsequently identified a second consanguineous family with SERKAL syndrome in which an affected fetus had CDH and an affected child had OFC. We then demonstrated that a subset of Wnt4 null embryos have perimembranous VSDs, anterior and posterior sac CDH, and soft palate clefts. These findings suggest that WNT4 deficiency can cause VSD, CDH, and palatal anomalies in mice and humans with SERKAL syndrome. These studies also suggest that our machine learning approach can be used as a candidate gene prioritization tool, and that targeted mouse phenotyping can serve as a means of confirming the roles of candidate genes in mammalian development.
Recent studies suggest that infection reprograms hematopoietic stem and progenitor cells (HSPCs) to enhance innate immune responses upon secondary infectious challenge, a process called "trained immunity."However, the specificity and cell types responsible for this response remain poorly defined. We established a model of trained immunity in mice in response to Mycobacterium avium infection. scRNA-seq analysis revealed that HSPCs activate interferon gamma-response genes heterogeneously upon primary challenge, while rare cell populations expand. Macrophages derived from trained HSPCs demonstrated enhanced bacterial killing and metabolism, and a single dose of recombinant interferon gamma exposure was sufficient to induce similar training. Mice transplanted with influenza -trained HSPCs displayed enhanced immunity against M. avium challenge and vice versa, demonstrating cross protection against antigenically distinct pathogens. Together, these results indicate that heterogeneous responses to infection by HSPCs can lead to long-term production of bone marrow derived macrophages with enhanced function and confer cross-protection against alternative pathogens.
We present the Causal Pivot (CP) as a structural causal model (SCM) for analyzing genetic heterogeneity in complex diseases. The CP leverages an established causal factor or factors to detect the contribution of additional suspected causes. Specifically, polygenic risk scores (PRSs) serve as known causes, while rare variants (RVs) or RV ensembles are evaluated as candidate causes. The CP incorporates outcome-induced association by conditioning on disease status. We derive a conditional maximum-likelihood procedure for binary and quantitative traits and develop the Causal Pivot likelihood ratio test (CP-LRT) to detect causal signals. Through simulations, we demonstrate the CP-LRT's robust power and superior error control compared to alternatives. We apply the CP-LRT to UK Biobank (UKB) data, analyzing three exemplar diseases: hypercholesterolemia (HC, low-density lipoprotein cholesterol ≥4.9 mmol/L; nc = 24,656), breast cancer (BC, ICD-10 C50; nc = 12,479), and Parkinson disease (PD, ICD-10 G20; nc = 2,940). For PRS, we utilize UKB-derived values, and for RVs, we analyze ClinVar pathogenic/likely pathogenic variants and loss-of-function mutations in disease-relevant genes: LDLR for HC, BRCA1 for BC, and GBA1 for PD. Significant CP-LRT signals were detected for all three diseases. Cross-disease and synonymous variant analyses serve as controls. We further develop ancestry adjustment using matching and inverse probability weighting as well as regression and doubly robust methods; we extend this to examine oligogenic burden in the lysosomal storage pathway in PD. The CP reveals an approach to address heterogeneity and is an extensible method for inference and discovery in complex disease genetics.
Background Disease-causing copy-number variants (CNVs) often encompass contiguous genes and can be detected using chromosomal microarray analysis (CMA). Conversely, CNVs affecting single disease-causing genes have historically been challenging to detect due to their small sizes. Methods A custom comprehensive CMA (Baylor College of Medicine - BCM v11.2) containing 400k probes and featuring exonic coverage for >4200 known or candidate disease-causing genes was utilized for the detection of CNVs at single-exon resolution. CMA results across a consecutive clinical cohort of more than 13 000 patients referred for genetic investigation at Baylor Genetics were examined. The genomic characteristics of CNVs impacting single protein-coding genes were investigated. Results Pathogenic or likely pathogenic (P/LP) CNVs (n = 190) affecting single protein-coding genes were detected in 188 patients, accounting for 9.9% (188/1894) of patients with P/LP CMA findings. The P/LP monogenic CNVs accounted for 9.2% (190/2058) of all P/LP nuclear CNVs detected by CMA. A total of 57.9% (110/190) of P/LP monogenic CNVs were smaller than 50 kb in size. Single exons were affected by 26.3% (50/190) of P/LP monogenic CNVs while 13.2% (25/190) affected 2 exons. CNVs were detected across 107 unique genes associated with predominantly autosomal dominant (AD) and X-linked (XL) conditions but also contributed to autosomal recessive (AR) conditions. Conclusions CMA with exon-targeted coverage of disease-associated genes facilitated the detection of small CNVs affecting single protein-coding genes, adding substantial clinical sensitivity to comprehensive CNV investigation. This approach resolved monogenic CNVs associated with autosomal and X-linked monogenic etiologies and yielded multiple significant findings. Monogenic CNVs represent an underrecognized subset of disease-causing alleles for Mendelian disorders.
Congenital Anomalies of Kidney and Urinary Tract (CAKUT) can occur in isolation or in conjunction with one or more non-CAKUT associated congenital anomalies or neurodevelopmental disorders (CAKUT+). A molecular cause is not identified in most individuals with CAKUT+. This is due, in part, to uncertainty regarding the efficacy of genetic testing and an incomplete understanding of the genes that cause CAKUT+. Here, we use data from 515 individuals with CAKUT+ (n = 500) or isolated CAKUT (n = 15) to determine the efficacy of clinical exome sequencing (cES) and to identify new phenotype expansions that involve CAKUT. We determined that cES established a molecular diagnosis in 27.4% (141/515) of individuals in this cohort. No statistically significant difference in efficacy was seen with regards to age, sex, CAKUT phenotype, or associated organ system abnormality. Only 3.5% (5/144) to 14.6% (21/144) of the individual diagnoses made in our cohort could have been identified using one of four clinically available CAKUT gene panels. We then used a machine-learning approach to confirm that PHIP is a CAKUT gene and to implicate ADNP and SETD5 genes associated with an increased risk of CAKUT. These findings lead us to conclude that cES should be considered in individuals with CAKUT+ for whom a molecular diagnosis has not been identified, that cES has the potential to identify many diagnoses in individuals with CAKUT+ that would be missed using a CAKUT gene panel, and that individuals with ADNP-, PHIP-, and SETD5-related disorders may present with CAKUT phenotypes.
Tetralogy of Fallot (TOF) is the most common cyanotic congenital heart defect (CHD). TOF may present in isolation or in conjunction with one or more non-cardiac congenital anomalies or neurodevelopmental disorders (TOF+). Uncertainty regarding the efficacy of various genetic testing strategies, and an incomplete understanding of the genetic causes of TOF+, may lead to hesitancy in recommending genetic testing, particularly, clinical exome sequencing (cES). Here, we analyzed cES data from 131 individuals with TOF+. A definitive or probable diagnosis was made for 31 individuals, yielding a diagnostic rate of 23.6% (31/131). One individual received three diagnoses. Commercially available CHD panels would have detected only 27.3% (9/33) to 63.6% (21/33) of the diagnoses made by cES. We then used a machine learning approach to identify four genes for which there is sufficient evidence to support a phenotypic expansion including TOF: DVL3, MED13L, PUF60, and MEIS2. Since chromosomal microarray analysis (CMA) has been reported to have a diagnostic efficacy of 10-20% in individuals with TOF, we conclude that cES should be considered for all individuals with TOF+ for whom a molecular diagnosis has not been established by CMA. We also conclude that TOF represents a low penetrance phenotype associated with genetic syndromes caused by pathogenic variants in DVL3, MED13L, PUF60, and MEIS2.
There is mounting evidence of the value of clinical genome sequencing (cGS) in individuals with suspected rare genetic disease (RGD), but cGS performance and impact on clinical care in a diverse population drawn from both high-income countries (HICs) and low- and middle-income countries (LMICs) has not been investigated. The iHope program, a philanthropic cGS initiative, established a network of 24 clinical sites in eight countries through which it provided cGS to individuals with signs or symptoms of an RGD and constrained access to molecular testing. A total of 1,004 individuals (median age, 6.5 years; 53.5% male) with diverse ancestral backgrounds (51.8% non-majority European) were assessed from June 2016 to September 2021. The diagnostic yield of cGS was 41.4% (416/1,004), with individuals from LMIC sites 1.7 times more likely to receive a positive test result compared to HIC sites (LMIC 56.5% [195/345] vs. HIC 33.5% [221/659], OR 2.6, 95% CI 1.9-3.4, p < 0.0001). A change in diagnostic evaluation occurred in 76.9% (514/668) of individuals. Change of management, inclusive of specialty referrals, imaging and testing, therapeutic interventions, and palliative care, was reported in 41.4% (285/694) of individuals, which increased to 69.2% (480/694) when genetic counseling and avoidance of additional testing were also included. Individuals from LMIC sites were as likely as their HIC counterparts to experience a change in diagnostic evaluation (OR 6.1, 95% CI 1.1-infinity, p = 0.05) and change of management (OR 0.9, 95% CI 0.5-1.3, p = 0.49). Increased access to genomic testing may support diagnostic equity and the reduction of global health care disparities.
RNA viruses, like SARS-CoV-2, depend on their RNA-dependent RNA polymerases (RdRp) for replication, which is error prone. Monitoring replication errors is crucial for understanding the virus's evolution. Current methods lack the precision to detect rare de novo RNA mutations, particularly in low-input samples such as those from patients. Here we introduce a targeted accurate RNA consensus sequencing method (tARC-seq) to accurately determine the mutation frequency and types in SARS-CoV-2, both in cell culture and clinical samples. Our findings show an average of 2.68 × 10−5 de novo errors per cycle with a C > T bias that cannot be solely attributed to APOBEC editing. We identified hotspots and cold spots throughout the genome, correlating with high or low GC content, and pinpointed transcription regulatory sites as regions more susceptible to errors. tARC-seq captured template switching events including insertions, deletions and complex mutations. These insights shed light on the genetic diversity generation and evolutionary dynamics of SARS-CoV-2. Targeted accurate RNA consensus sequencing enables study of de novo errors caused by RNA-dependent RNA polymerases and provides deeper insights into how SARS-CoV-2 genetic diversity emerges.
FOXP1 encodes a transcription factor involved in tissue regulation and cell-type-specific functions. Haploinsufficiency of FOXP1 is associated with a neurodevelopmental disorder: autosomal dominant mental retardation with language impairment with or without autistic features. More recently, heterozygous FOXP1 variants have also been shown to cause a variety of structural birth defects including central nervous system (CNS) anomalies, congenital heart defects, congenital anomalies of the kidney and urinary tract, cryptorchidism, and hypospadias. In this report, we present a previously unpublished case of an individual with congenital diaphragmatic hernia (CDH) who carries an approximately 3.8 Mb deletion. Based on this deletion, and deletions previously reported in two other individuals with CDH, we define a CDH critical region on chromosome 3p13 that includes FOXP1 and four other protein-coding genes. We also provide detailed clinical descriptions of two previously reported individuals with CDH who carry de novo, pathogenic variants in FOXP1 that are predicted to trigger nonsense-mediated mRNA decay. A subset of individuals with putatively deleterious FOXP4 variants has also been shown to develop CDH. Since FOXP proteins function as homo- or heterodimers and the homologs of FOXP1 and FOXP4 are expressed at the same time points in the embryonic mouse diaphragm, they may function together as a dimer, or in parallel as homodimers, to regulate gene expression during diaphragm development. Not all individuals with heterozygous, loss-of-function changes in FOXP1 develop CDH. Hence, we conclude that FOXP1 acts as a susceptibility factor that contributes to the development of CDH in conjunction with other genetic, epigenetic, environmental, and/or stochastic factors.
PURPOSE. A molecular diagnosis is only made in a subset of individuals with nonisolated microphthalmia, anophthalmia, and coloboma (MAC). This may be due to underutilization of clinical (whole) exome sequencing (cES) and an incomplete understanding of the genes that cause MAC. The purpose of this study is to determine the efficacy of cES in cases of nonisolated MAC and to identify new MAC phenotypic expansions. METHODS. We determined the efficacy of cES in 189 individuals with nonisolated MAC. We then used cES data, a validated machine learning algorithm, and previously published expression data, case reports, and animal models to determine which candidate genes were most likely to contribute to the development of MAC. RESULTS. We found the efficacy of cES in nonisolated MAC to be between 32.3% (61/189) and 48.1% (91/189). Most genes affected in our cohort were not among genes currently screened in clinically available ophthalmologic gene panels. A subset of the genes implicated in our cohort had not been clearly associated with MAC. Our analyses revealed sufficient evidence to support low-penetrance MAC phenotypic expansions involving nine of these human disease genes. CONCLUSIONS. We conclude that cES is an effective means of identifying a molecular diagnosis in individuals with nonisolated MAC and may identify putatively damaging variants that would be missed if only a clinically available ophthalmologic gene panel was obtained. Our data also suggest that deleterious variants in BRCA2, BRIP1, KAT6A, KAT6B, NSF, RAC1, SMARCA4, SMC1A, and TUBA1A can contribute to the development of MAC.
ABSTRACT Purpose The variome of the Turkish (TK) population, a population with a considerable history of admixture and consanguinity, has not been deeply investigated deeply for its potential impact on the genomic architecture of disease traits. Methods We generated and analyzed a database of variants derived from exome sequencing (ES) data of 773 TK unrelated, clinically affected individuals with various suspected Mendelian disease traits, and 643 unaffected relatives. Results Using uniform manifold approximation and projection (UMAP), we showed that the TK genomes are more similar to those of Europeans and consist of two main subpopulations: clusters 1 and 2 (N=235 and 1,181) that differ in admixture proportion and variome ( https://turkishvariomedb.shinyapps.io/tvdb/ ). Furthermore, the higher inbreeding coefficient ( F ) values observed in the TK affected compared to unaffected individuals correlated with a larger median span of long-sized (>2.64 Mb) runs of homozygosity (ROH) regions ( p -value=2.09e-18). We show that long-sized ROHs are more likely to be formed on recently configured haplotypes enriched for rare homozygous deleterious variants in the TK-affected compared to TK-unaffected individuals ( p -value= 3.35e-11). Analysis of genotype-phenotype correlations reveals that genes with rare homozygous deleterious variants in long-sized ROHs provide the most comprehensive set of molecular diagnoses for the observed disease traits with a systematic quantitative analysis of HPO (Human Phenotype Ontology) terms. Conclusion Our findings support the notion that novel rare variants on newly configured haplotypes arising within the recent past generations of a family or clan contribute significantly to recessive disease traits in the TK population.
Trained immunity is defined as long-term epigenetic and metabolic reprogramming of innate immune cells resulting in an augmented, non-specific response to a secondary stimulus. We have shown that HSPCs can store long-term memory, and macrophages derived from these trained HSPCs demonstrate increased metabolism, cytokine production, and bacterial killing. However, the specific HSPC subpopulation that harbors trained immunity remains elusive. Our single-cell RNA seq studies uncovered an emerging population, named infection-activated-hematopoietic stem cells (IA-HSCs) that were elicited by Mycobacterium avium (M. avium) and were found to have stem cell characteristics. To further investigate whether IA-HSCs are critical for HSPC-trained immunity. We tested numerous infectious stimuli including M. avium, Bacillus Calmette-Guerin vaccine, lipopolysaccharide (LPS), and influenza H1N1 in eliciting IA-HSCs (Lin-cKit+Sca1+CD24a+CD81+CD9+). We found that IA-HSCs were most effectively induced by M. avium infection. We transplanted 3000 IA-HSCs induced by M. avium or 200000 M. avium trained or untrained cKit+ HSPCs into lethally irradiated C57BL/6 CD45.1 recipient mice. 16 weeks post-transplant, CD45.2 HSPCs were isolated from the recipient bone marrow and differentiated into bone marrow-derived macrophages (BMDMs) in vitro. Hallmarks of trained immunity phenotype were assessed through in vitro killing assay, seahorse assays, and cytokine production. We found that BMDMs derived from IA-HSCs had significantly lower M. avium burden. Furthermore, they also had enhanced metabolic function, and elevated IL-6 and TNF-α production compared to the untrained and the trained cKit+ HSPC groups. Overall, we conclude that IA-HSCs alone conferred trained immunity to downstream BMDMs, demonstrating that these cells are transplantable and capable of encoding long-term innate immune memory.
Abstract Breast cancer research is hampered by difficulties in obtaining and studying primary human breast tissue, and by the lack of in vivo preclinical models that reflect patient tumor biology accurately. To overcome these limitations, we propagated a cohort of human breast tumors grown in the epithelium-free mammary fat pad of severe combined immunodeficient (SCID)/Beige and nonobese diabetic (NOD)/SCID/IL-2γ-receptor null (NSG) mice under a series of transplant conditions. Both models yielded stably transplantable xenografts at comparably high rates (∼21% and ∼19%, respectively). Of the conditions tested, xenograft take rate was highest in the presence of a low-dose estradiol pellet. Overall, 32 stably transplantable xenograft lines were established, representing 25 unique patients. Most tumors yielding xenografts were “triple-negative” [estrogen receptor (ER)−progesterone receptor (PR)−HER2+; n = 19]. However, we established lines from 3 ER−PR−HER2+ tumors, one ER+PR−HER2−, one ER+PR+HER2−, and one “triple-positive” (ER+PR+HER2+) tumor. Serially passaged xenografts show biologic consistency with the tumor of origin, are phenotypically stable across multiple transplant generations at the histologic, transcriptomic, proteomic, and genomic levels, and show comparable treatment responses as those observed clinically. Xenografts representing 12 patients, including 2 ER+ lines, showed metastasis to the mouse lung. These models thus serve as a renewable, quality-controlled tissue resource for preclinical studies investigating treatment response and metastasis. Cancer Res; 73(15); 4885–97. ©2013 AACR.
PURPOSE:Genome sequencing (GS) may shorten the diagnostic odyssey for patients, but clinical experience with this assay in nonresearch settings remains limited. Texas Children's Hospital began offering GS as a clinical test to admitted patients in 2020, providing an opportunity to study GS utilization, possibilities for test optimization, and testing outcomes.METHODS:We retrospectively reviewed GS orders for admitted patients for a nearly 3-year period from March 2020 through December 2022. We gathered anonymized clinical data from the electronic health record to answer the study questions.RESULTS:The diagnostic yield over 97 admitted patients was 35%. The majority of GS clinical indications were neurologic or metabolic (61%) and most patients were in intensive care (58%). Tests were often characterized as candidates for intervention/improvement (56%), frequently because of redundancy with prior testing. Patients receiving GS without prior exome sequencing (ES) had higher diagnostic rates (45%) than the cohort as a whole. In 2 cases, GS revealed a molecular diagnosis that is unlikely to be detected by ES.CONCLUSION:The performance of GS in clinical settings likely justifies its use as a first-line diagnostic test, but the incremental benefit for patients with prior ES may be limited.
Evidence suggests that genetic factors contribute to the development of anorectal malformations (ARMs). However, the etiology of the majority of ARMs cases remains unclear. Exome sequencing (ES) may be underutilized in the diagnostic workup of ARMs due to uncertainty regarding its diagnostic yield. In a clinical database of ~17,000 individuals referred for ES, we identified 130 individuals with syndromic ARMs. A definitive or probable diagnosis was made in 45 of these individuals for a diagnostic yield of 34.6% (45/130). The molecular diagnostic yield of individuals who initially met criteria for VACTERL association was lower than those who did not (26.8% vs 44.1%; p = 0.0437), suggesting that non-genetic factors may play an important role in this subset of syndromic ARM cases. Within this cohort, we identified two individuals who carried de novo pathogenic frameshift variants in ADNP, two individuals who were homozygous for pathogenic variants in BBS1, and single individuals who carried pathogenic or likely pathogenic variants in CREBBP, EP300, FANCC, KDM6A, SETD2, and SMARCA4. The association of these genes with ARMs was supported by previously published cases, and their similarity to known ARM genes as demonstrated using a machine learning algorithm. These data suggest that ES should be considered for all individuals with syndromic ARMs in whom a molecular diagnosis has not been made, and that ARMs represent a low penetrance phenotype associated with Helsmoortel-van der Aa syndrome, Bardet-Biedl syndrome 1, Rubinstein-Taybi syndromes 1 and 2, Fanconi anemia group C, Kabuki syndrome 2, SETD2-related disorders, and Coffin-Siris syndrome 4.