This study was aimed at assessing the diagnostic utility of whole genome sequence analysis in a well-characterised research cohort of individuals referred with a clinical suspicion of Cornelia de Lange syndrome (CdLS) in whom prior genetic testing had not identified a causative variant. Short-read whole genome sequencing was performed on 195 individuals from 105 families, 108 of whom were affected. 100/108 of the affected individuals had prior relevant genetic testing, with no pathogenic variant being identified. The study group comprised 42 trios in which both parental samples were available for testing (42 affected individuals and 126 unaffected parents), 61 singletons (unrelated affected individuals), and two families with more than one affected individual. The results showed that 32 unrelated probands from 105 families (30.5%) had likely causative coding region-disrupting variants. Four loci were identified in > 1 proband: NIPBL (10), ANKRD11 (6), EP300 (3), and EHMT1 (2). Single variants were detected in the remaining genes (EBF3, KMT2A, MED13L, NLGN3, NR2F1, PHIP, PUF60, SET, SETD5, SMC1A, and TBL1XR1). Possibly causative variants in noncoding regions of NIPBL were identified in four individuals. Single de novo variants were identified in five genes not previously reported to be associated with any developmental disorder: ARID3A, PIK3C3, MCM7, MIS18BP1, and WDR18. The clustering of de novo noncoding variants implicates a single upstream open reading frame (uORF) and a small region in Intron 21 in NIPBL regulation. Causative variants in genes encoding chromatin-associated proteins, with no defined influence on cohesin function, appear to result in CdLS-like clinical features. This study demonstrates the clinical utility of whole genome sequencing as a diagnostic test in individuals presenting with CdLS or CdLS-like phenotypes.
Heterozygous missense variants and in-frame indels in SMC3 are a cause of Cornelia de Lange syndrome (CdLS), marked by intellectual disability, growth deficiency, and dysmorphism, via an apparent dominant-negative mechanism. However, the spectrum of manifestations associated with SMC3 loss-of-function variants has not been reported, leading to hypotheses of alternative phenotypes or even developmental lethality. We used matchmaking servers, patient registries, and other resources to identify individuals with heterozygous, predicted loss-of-function (pLoF) variants in SMC3, and analyzed population databases to characterize mutational intolerance in this gene. Here, we show that SMC3 behaves as an archetypal haploinsufficient gene: it is highly constrained against pLoF variants, strongly depleted for missense variants, and pLoF variants are associated with a range of developmental phenotypes. Among 14 individuals with SMC3 pLoF variants, phenotypes were variable but coalesced on low growth parameters, developmental delay/intellectual disability, and dysmorphism, reminiscent of atypical CdLS. Comparisons to individuals with SMC3 missense/in-frame indel variants demonstrated an overall milder presentation in pLoF carriers. Furthermore, several individuals harboring pLoF variants in SMC3 were nonpenetrant for growth, developmental, and/or dysmorphic features, and some had alternative symptomatologies with rational biological links to SMC3. Analyses of tumor and model system transcriptomic data and epigenetic data in a subset of cases suggest that SMC3 pLoF variants reduce SMC3 expression but do not strongly support clustering with functional genomic signatures of typical CdLS. Our finding of substantial population-scale LoF intolerance in concert with variable growth and developmental features in subjects with SMC3 pLoF variants expands the scope of cohesinopathies, informs on their allelic architecture, and suggests the existence of additional clearly LoF-constrained genes whose disease links will be confirmed only by multi-layered genomic data paired with careful phenotyping.
Background Classic aniridia is a highly penetrant autosomal dominant disorder characterised by congenital absence of the iris, foveal hypoplasia, optic disc anomalies and progressive opacification of the cornea. >90% of cases of classic aniridia are caused by heterozygous, loss-of-function variants affecting the PAX6 locus.Methods Short-read whole genome sequencing was performed on 51 (39 affected) individuals from 37 different families who had screened negative for mutations in the PAX6 coding region.Results Likely causative mutations were identified in 22 out of 37 (59%) families. In 19 out of 22 families, the causative genomic changes have an interpretable deleterious impact on the PAX6 locus. Of these 19 families, 1 has a novel heterozygous PAX6 frameshift variant missed on previous screens, 4 have single nucleotide variants (SNVs) (one novel) affecting essential splice sites of PAX6 5 ' non-coding exons and 2 have deep intronic SNV (one novel) resulting in gain of a donor splice site. In 12 out of 19, the causative variants are large-scale structural variants; 5 have partial or whole gene deletions of PAX6, 3 have deletions encompassing critical PAX6 cis-regulatory elements, 2 have balanced inversions with disruptive breakpoints within the PAX6 locus and 2 have complex rearrangements disrupting PAX6. The remaining 3 of 22 families have deletions encompassing FOXC1 (a known cause of atypical aniridia). Seven of the causative variants occurred de novo and one cosegregated with familial aniridia. We were unable to establish inheritance status in the remaining probands. No plausibly causative SNVs were identified in PAX6 cis-regulatory elements.Conclusion Whole genome sequencing proves to be an effective diagnostic test in most individuals with previously unexplained aniridia.
The emergence of retinal progenitor cells and differentiation to various retinal cell types represent fundamental processes during retinal development. Herein, we provide a comprehensive single cell characterisation of transcriptional and chromatin accessibility changes that underline retinal progenitor cell specification and differentiation over the course of human retinal development up to midgestation. Our lineage trajectory data demonstrate the presence of early retinal progenitors, which transit to late, and further to transient neurogenic progenitors, that give rise to all the retinal neurons. Combining single cell RNA-Seq with spatial transcriptomics of early eye samples, we demonstrate the transient presence of early retinal progenitors in the ciliary margin zone with decreasing occurrence from 8 post-conception week of human development. In retinal progenitor cells, we identified a significant enrichment for transcriptional enhanced associate domain transcription factor binding motifs, which when inhibited led to loss of cycling progenitors and retinal identity in pluripotent stem cell derived organoids. Formation of the retina during development involves the coordinated action of retinal progenitor cells and their differentiated cell types, which is key for producing a functioning eye. Here the authors provide a detailed atlas of human retinal development, combining scRNA-seq and spatial transcriptomics, and identify key genetic factors that mediate retinal progenitor cell proliferation and differentiation.
Abstract Nonsense and missense mutations in the transcription factor PAX6 cause a wide range of eye development defects, including aniridia, microphthalmia and coloboma. To understand how changes of PAX6:DNA binding cause these phenotypes, we combined saturation mutagenesis of the paired domain of PAX6 with a yeast one-hybrid (Y1H) assay in which expression of a PAX6-GAL4 fusion gene drives antibiotic resistance. We quantified binding of more than 2700 single amino-acid variants to two DNA sequence elements. Mutations in DNA-facing residues of the N-terminal subdomain and linker region were most detrimental, as were mutations to prolines and to negatively charged residues. Many variants caused sequence-specific molecular gain-of-function effects, including variants in position 71 that increased binding to the LE9 enhancer but decreased binding to a SELEX-derived binding site. In the absence of antibiotic selection, variants that retained DNA binding slowed yeast growth, likely because such variants perturbed the yeast transcriptome. Benchmarking against known patient variants and applying ACMG/AMP guidelines to variant classification, we obtained supporting-to-moderate evidence that 977 variants are likely pathogenic and 1306 are likely benign. Our analysis shows that most pathogenic mutations in the paired domain of PAX6 can be explained simply by the effects of these mutations on PAX6:DNA association, and establishes Y1H as a generalisable assay for the interpretation of variant effects in transcription factors.
Purpose Structural mosaicism has been previously implicated in developmental disorders. We aimed to identify rare mosaic chromosomal alterations (MCAs) in probands with severe undiagnosed developmental disorders. Methods We identified MCAs in genotyping array data from 12,530 probands in the Deciphering Developmental Disorders study using mosaic chromosome alterations caller (MoChA). Results We found 61 MCAs in 57 probands, many of these were tissue specific. In 23 of 26 (88.5%) cases for which the MCA was detected in saliva in which blood was also available for analysis, the MCA could not be detected in blood. The MCAs included 20 polysomies, comprising either 1 arm of a chromosome or a whole chromosome, for which we were able to show the timing of the error (25% mitosis, 40% meiosis I, and 35% meiosis II). Only 2 of 57 (3.5%) of the probands in whom we found MCAs had another likely genetic diagnosis identified by exome sequencing, despite an overall diagnostic yield of ∼40% across the cohort. Conclusion Our results show that identification of MCAs provides candidate diagnoses for previously undiagnosed patients with developmental disorders, potentially explaining ∼0.45% of cases in the Deciphering Developmental Disorders study. Nearly 90% of these MCAs would have remained undetected by analyzing DNA from blood and no other tissue.
Specification of the eye field (EF) within the neural plate marks the earliest detectable stage of eye development. Experimental evidence, primarily from non-mammalian model systems, indicates that the stable formation of this group of cells requires the activation of a set of key transcription factors. This crucial event is challenging to probe in mammals and, quantitatively, little is known regarding the regulation of the transition of cells to this ocular fate. Using optic vesicle organoids to model the onset of the EF, we generate time-course transcriptomic data allowing us to identify dynamic gene expression programmes that characterize this cellular-state transition. Integrating this with chromatin accessibility data suggests a direct role of canonical EF transcription factors in regulating these gene expression changes, and highlights candidate cis-regulatory elements through which these transcription factors act. Finally, we begin to test a subset of these candidate enhancer elements, within the organoid system, by perturbing the underlying DNA sequence and measuring transcriptomic changes during EF activation.
Purpose: Transformer2 proteins (Tra2a and Tra2(3) control splicing patterns in human cells, and no human phenotypes have been associated with germline variants in these genes. The aim of this work was to associate germline variants in the TRA2B gene to a novel neurodevelopmental disorder.Methods: A total of 12 individuals from 11 unrelated families who harbored predicted loss-of -function monoallelic variants, mostly de novo, were recruited. RNA sequencing and western blot analyses of Tra2(3-1 and Tra2(3-3 isoforms from patient-derived cells were performed. Tra2(31-GFP, Tra2(33-GFP and CHEK1 exon 3 plasmids were transfected into HEK-293 cells.Results: All variants clustered in the 5 ' part of TRA2B, upstream of an alternative translation start site responsible for the expression of the noncanonical Tra2(3-3 isoform. All affected in-dividuals presented intellectual disability and/or developmental delay, frequently associated with infantile spasms, microcephaly, brain anomalies, autism spectrum disorder, feeding difficulties, and short stature. Experimental studies showed that these variants decreased the expression of the canonical Tra2(3-1 isoform, whereas they increased the expression of the Tra2(3-3 isoform, which is shorter and lacks the N-terminal RS1 domain. Increased expression of Tra2(3-3-GFP were shown to interfere with the incorporation of CHEK1 exon 3 into its mature transcript, normally incorporated by Tra2(3-1.Conclusion: Predicted loss-of-function variants clustered in the 5 ' portion of TRA2B cause a new neurodevelopmental syndrome through an apparently dominant negative disease mechanism involving the use of an alternative translation start site and the overexpression of a shorter, repressive Tra2(3 protein.(c) 2022 The Authors. Published by Elsevier Inc. on behalf of American College of Medical Genetics and Genomics. This is an open access article under the CC BY license (http://creativecommons.org/licenses/by/4.0/).
BACKGROUND:Genomic variant prioritisation is one of the most significant bottlenecks to mainstream genomic testing in healthcare. Tools to improve precision while ensuring high recall are critical to successful mainstream clinical genomic testing, in particular for whole genome sequencing where millions of variants must be considered for each patient. METHODS:We developed EyeG2P, a publicly available database and web application using the Ensembl Variant Effect Predictor. EyeG2P is tailored for efficient variant prioritisation for individuals with inherited ophthalmic conditions. We assessed the sensitivity of EyeG2P in 1234 individuals with a broad range of eye conditions who had previously received a confirmed molecular diagnosis through routine genomic diagnostic approaches. For a prospective cohort of 83 individuals, we assessed the precision of EyeG2P in comparison with routine diagnostic approaches. For 10 additional individuals, we assessed the utility of EyeG2P for whole genome analysis. RESULTS:EyeG2P had 99.5% sensitivity for genomic variants previously identified as clinically relevant through routine diagnostic analysis (n=1234 individuals). Prospectively, EyeG2P enabled a significant increase in precision (35% on average) in comparison with routine testing strategies (p<0.001). We demonstrate that incorporation of EyeG2P into whole genome sequencing analysis strategies can reduce the number of variants for analysis to six variants, on average, while maintaining high diagnostic yield. CONCLUSION:Automated filtering of genomic variants through EyeG2P can increase the efficiency of diagnostic testing for individuals with a broad range of inherited ophthalmic disorders.
BackgroundPediatric disorders include a range of highly penetrant, genetically heterogeneous conditions amenable to genomewide diagnostic approaches. Finding a molecular diagnosis is challenging but can have profound lifelong benefits.MethodsWe conducted a large-scale sequencing study involving more than 13,500 families with probands with severe, probably monogenic, difficult-to-diagnose developmental disorders from 24 regional genetics services in the United Kingdom and Ireland. Standardized phenotypic data were collected, and exome sequencing and microarray analyses were performed to investigate novel genetic causes. We developed an iterative variant analysis pipeline and reported candidate variants to clinical teams for validation and diagnostic interpretation to inform communication with families. Multiple regression analyses were performed to evaluate factors affecting the probability of diagnosis.ResultsA total of 13,449 probands were included in the analyses. On average, we reported 1.0 candidate variant per parent-offspring trio and 2.5 variants per singleton proband. With the use of clinical and computational approaches to variant classification, a diagnosis was made in approximately 41% of probands (5502 of 13,449), of whom 76% had a pathogenic de novo variant. Another 22% of probands (2997 of 13,449) had variants of uncertain significance in genes that were strongly linked to monogenic developmental disorders. Recruitment in a parent-offspring trio had the largest effect on the probability of diagnosis (odds ratio, 4.70; 95% confidence interval [CI], 4.16 to 5.31). Probands were less likely to receive a diagnosis if they were born extremely prematurely (i.e., 22 to 27 weeks' gestation; odds ratio, 0.39; 95% CI, 0.22 to 0.68), had in utero exposure to antiepileptic medications (odds ratio, 0.44; 95% CI, 0.29 to 0.67), had mothers with diabetes (odds ratio, 0.52; 95% CI, 0.41 to 0.67), or were of African ancestry (odds ratio, 0.51; 95% CI, 0.31 to 0.78).ConclusionsAmong probands with severe, probably monogenic, difficult-to-diagnose developmental disorders, multimodal analysis of genomewide data had good diagnostic power, even after previous attempts at diagnosis. (Funded by the Health Innovation Challenge Fund and Wellcome Sanger Institute.)
Diagnosing rare developmental disorders using genome-wide sequencing data commonly necessitates review of multiple plausible candidate variants, often using ontologies of categorical clinical terms. We show that Integrating Multiple Phenotype Resources Optimizes Variant Evaluation in Developmental Disorders (IMPROVE-DD) by incorporating additional classes of data commonly available to clinicians and recorded in health records. In doing so, we quantify the distinct contributions of sex, growth, and development in addition to Human Phenotype Ontology (HPO) terms and demonstrate added value from these readily available information sources. We use likelihood ratios for nominal and quantitative data and propose a classifier for HPO terms in this framework. This Bayesian framework results in more robust diagnoses. Using data systematically collected in the Deciphering Developmental Disorders study, we considered 77 genes with pathogenic/likely pathogenic variants in ≥10 individuals. All genes showed at least a satisfactory prediction by receiver operating characteristic when testing on training data (AUC ≥ 0.6), and HPO terms were the best predictor for the majority of genes, though a minority (13/77) of genes were better predicted by other phenotypic data types. Overall, classifiers based upon multiple integrated phenotypic data sources performed better than those based upon any individual source, and importantly, integrated models produced notably fewer false positives. Finally, we show that IMPROVE-DD models with good predictive performance on cross-validation can be constructed from relatively few individuals. This suggests new strategies for candidate gene prioritization and highlights the value of systematic clinical data collection to support diagnostic programs.
AbstractHeterozygous missense variants and in-frame indels inSMC3are a cause of Cornelia de Lange syndrome (CdLS), marked by intellectual disability, growth deficiency, and dysmorphism, via an apparent dominant-negative mechanism. However, the spectrum of manifestations associated withSMC3loss-of-function variants has not been reported, leading to hypotheses of alternative phenotypes or even developmental lethality. We used matchmaking servers, patient registries, and other resources to identify individuals with heterozygous, predicted loss-of-function (pLoF) variants inSMC3, and analyzed population databases to characterize mutational intolerance in this gene. Here, we show thatSMC3behaves as an archetypal haploinsufficient gene: it is highly constrained against pLoF variants, strongly depleted for missense variants, and pLoF variants are associated with a range of developmental phenotypes. Among 13 individuals withSMC3pLoF variants, phenotypes were variable but coalesced on low growth parameters, developmental delay/intellectual disability, and dysmorphism reminiscent of atypical CdLS. Comparisons to individuals withSMC3missense/in-frame indel variants demonstrated a milder presentation in pLoF carriers. Furthermore, several individuals harboring pLoF variants inSMC3were nonpenetrant for growth, developmental, and/or dysmorphic features, some instead having intriguing symptomatologies with rational biological links toSMC3including bone marrow failure, acute myeloid leukemia, and Coats retinal vasculopathy. Analyses of transcriptomic and epigenetic data suggest thatSMC3pLoF variants reduceSMC3expression but do not result in a blood DNA methylation signature clustering with that of CdLS, and that the global transcriptional signature ofSMC3loss is model-dependent. Our finding of substantial population-scale LoF intolerance in concert with variable penetrance in subjects withSMC3pLoF variants expands the scope of cohesinopathies, informs on their allelic architecture, and suggests the existence of additional clearly LoF-constrained genes whose disease links will be confirmed only by multi-layered genomic data paired with careful phenotyping.
Human limbs emerge during the fourth post-conception week as mesenchymal buds, which develop into fully formed limbs over the subsequent months 1 . This process is orchestrated by numerous temporally and spatially restricted gene expression programmes, making congenital alterations in phenotype common 2 . Decades of work with model organisms have defined the fundamental mechanisms underlying vertebrate limb development, but an in-depth characterization of this process in humans has yet to be performed. Here we detail human embryonic limb development across space and time using single-cell and spatial transcriptomics. We demonstrate extensive diversification of cells from a few multipotent progenitors to myriad differentiated cell states, including several novel cell populations. We uncover two waves of human muscle development, each characterized by different cell states regulated by separate gene expression programmes, and identify musculin (MSC) as a key transcriptional repressor maintaining muscle stem cell identity. Through assembly of multiple anatomically continuous spatial transcriptomic samples using VisiumStitcher, we map cells across a sagittal section of a whole fetal hindlimb. We reveal a clear anatomical segregation between genes linked to brachydactyly and polysyndactyly, and uncover transcriptionally and spatially distinct populations of the mesenchyme in the autopod. Finally, we perform single-cell RNA sequencing on mouse embryonic limbs to facilitate cross-species developmental comparison, finding substantial homology between the two species.
Neurodevelopmental disorder with visual defects and brain anomalies (NEDVIBA) is a recently described genetic condition caused by de novo missense HK1 variants. Phenotypic data is currently limited; only seven patients have been published to date. This descriptive case series of a further four patients with de novo missense HK1 variants, alongside integration of phenotypic data with the reported cases, aims to improve our understanding of the associated phenotype. We provide further evidence that de novo HK1 variants located within the regulatory-terminal domain and alpha helix are associated with neurological problems and visual problems. We highlight for the first time an association with a raised cerebrospinal fluid lactate and specific abnormalities to the basal ganglia on brain magnetic resonance imaging, as well as associated respiratory issues and swallowing/feeding difficulties. We propose that this distinctive neurodevelopmental phenotype could arise through disruption of the regulatory glucose-6-phosphate binding site and subsequent gain of function of HK1 within the brain.
ABSTRACTBackgroundPediatric disorders include a range of highly genetically heterogeneous conditions that are amenable to genome-wide diagnostic approaches. Finding a molecular diagnosis is challenging but can have profound lifelong benefits.MethodsThe Deciphering Developmental Disorders (DDD) study recruited >33,500 individuals from families with severe, likely monogenic developmental disorders from 24 regional genetics services around the UK and Ireland. We collected detailed standardised phenotype data and performed whole-exome sequencing and microarray analysis to investigate novel genetic causes. We developed an augmented variant analysis and re-analysis pipeline to maximise sensitivity and specificity, and communicated candidate variants to clinical teams for validation and diagnostic interpretation. We performed multiple regression analyses to evaluate factors affecting the probability of being diagnosed.ResultsWe reported approximately one candidate variant per parent-offspring trio and 2.5 variants per singleton proband, including both sequence and structural variants. Using clinical and computational approaches to variant classification, we have achieved a diagnosis in at least 34% (4507 probands), of whom 67% have a pathogenicde novomutation. Being recruited as a parent-offspring trio had the largest impact on the chance of being diagnosed (OR=4.70). Probands who were extremely premature (OR=0.39), hadin uteroexposure to antiepileptic medications (OR=0.44), or whose mothers had diabetes (OR=0.52) were less likely to be diagnosed, as were those of African ancestry (OR=0.51).ConclusionsOptimising diagnosis and discovery in highly penetrant genomic disease depends upon ongoing and novel scientific analyses, ethical recruitment and feedback policies, and collaborative clinical-research partnerships.
Background The majority of clinical genetic testing focuses almost exclusively on regions of the genome that directly encode proteins. The important role of variants in non-coding regions in penetrant disease is, however, increasingly being demonstrated, and the use of whole genome sequencing in clinical diagnostic settings is rising across a large range of genetic disorders. Despite this, there is no existing guidance on how current guidelines designed primarily for variants in protein-coding regions should be adapted for variants identified in other genomic contexts. Methods We convened a panel of nine clinical and research scientists with wide-ranging expertise in clinical variant interpretation, with specific experience in variants within non-coding regions. This panel discussed and refined an initial draft of the guidelines which were then extensively tested and reviewed by external groups. Results We discuss considerations specifically for variants in non-coding regions of the genome. We outline how to define candidate regulatory elements, highlight examples of mechanisms through which non-coding region variants can lead to penetrant monogenic disease, and outline how existing guidelines can be adapted for the interpretation of these variants. Conclusions These recommendations aim to increase the number and range of non-coding region variants that can be clinically interpreted, which, together with a compatible phenotype, can lead to new diagnoses and catalyse the discovery of novel disease mechanisms.
PURPOSE Several groups and resources provide information that pertains to the validity of gene-disease relationships used in genomic medicine and research; however, universal standards and terminologies to define the evidence base for the role of a gene in disease, and a single harmonized resource were lacking. To tackle this issue, the Gene Curation Coalition (GenCC) was formed. METHODS The GenCC drafted harmonized definitions for differing levels of gene-disease validity based on existing resources, and performed a modified Delphi survey with three rounds to narrow the list of terms. The GenCC also developed a unified database to display curated gene-disease validity assertions from its members. RESULTS Based on 241 survey responses from the genetics community, a consensus term set was chosen for grading gene-disease validity and database submissions. As of December 2021, the database contains 15,241 gene-disease assertions on 4,569 unique genes from 12 submitters. When comparing submissions to the database from distinct sources, conflicts in assertions of gene-disease validity ranged from 5.3% to 13.4%. CONCLUSION Terminology standardization, sharing of gene-disease validity classifications, and resolution of curation conflicts will facilitate collaborations across international curation efforts and in turn, improve consistency in genetic testing and variant interpretation.