Balanced Robertsonian translocation (ROB) is the most common chromosomal rearrangement in humans, with an estimated occurrence of 1 in 800 in newborn studies. Carriers are at increased risk of cancer and often diagnosed at fertility clinics after facing recurrent miscarriages, infertility, or aneuploid offspring. Genotyping carriers with DNA sequencing has been challenging because of gaps and misrepresentation of the translocation fusion site in the human reference genome. Only recently, telomere-to-telomere (T2T) human genomes successfully revealed sequences of the acrocentric short arms, including the most common ROB fusion site. A ROB results in loss of two ribosomal DNA (rDNA) arrays and its adjacent distal sequences, including the highly conserved distal junction (DJ). Here, we present a novel method to type ROB carriers directly from short sequencing reads by estimating DJ copy number. We demonstrate that our method successfully genotypes ROBs using a reference-free approach or alignments to either T2T-CHM13v2 or GRCh38. Applying the method to a cohort of healthy newborns and family members (n=4,172) as well as the UK Biobank (n=490,416), we find candidate ROBs at a frequency consistent with the previously reported 1 in 800 incidence (0.11-0.12%). In addition to ROB carriers, we report the frequency of one DJ loss (9, 2.8-3.4%) or gain (11+, 8.4-9.3%) from the two cohorts and the 1000 Genomes Project (n=3,202), and characterize the underlying structural variation in near-T2T genome assemblies from the Human Pangenome Reference Consortium. Importantly, our method provides the first sequencing-based diagnostic for Robertsonian chromosomes and can be applied to low-coverage sequencing data, enhancing its clinical applicability and enabling new studies of structural variation on the acrocentric chromosomes.
The polychaete Capitella teleta is a primary model for evolutionary developmental biology, comparative genomics, conservation, and ecotoxicology. Although it was the first polychaete genome sequenced, the original assembly is outdated by modern standards. Here, we combine long-read and short-read sequencing with Hi-C chromatin conformation capture to assemble chromosome-level nuclear and mitochondrial genomes of the laboratory strain of C. teleta. This reference assembly accurately reflects the expected genome size (∼243.6 Mb) and contains a highly complete, evolutionarily conserved gene repertoire. Notably, the nuclear and mitochondrial genomes are heavily rearranged, indicating a decoupling between gene family repertoire and chromosomal evolution. The analyses of developmental time courses of bulk and single-cell RNA-seq, ATAC-seq, and EM-seq data using the new reference assembly resulted in a significant improvement in quality, enabling the identification of new cell-type-specific gene markers. Finally, we generated a publicly available genome browser that ensures these resources comply with FAIR principles. Our study provides state-of-the-art genomic resources for C. teleta, addressing a pressing community need and opening new research opportunities in animal and genome evolution.
Motivation:The colonial hydroid Hydractinia exhibits several unique biological properties, including its remarkable regenerative capacity and the ability to distinguish self from non-self, characteristics that make them valuable models for studying human disease and aging. The availability of well-annotated multi-omic data, as well as tools to visualize these data, is essential for advancing the use of these model organisms to enhance our understanding of the relationship between genomic and morphological complexity, the evolution of multicellularity, and the emergence of novel cell types. Results:We present the Hydractinia Genome Project Portal, a comprehensive resource providing genomic, transcriptomic, and proteomic datasets for two widely studied Hydractinia species. The portal provides extensive sequence, structure, and functional annotation resources that are not available elsewhere, including genome browsers, a single-cell gene expression atlas, a protein structure viewer, and a custom BLAST implementation. We demonstrate the portal's utility for biological discovery and have used a subset of Hydractinia-specific stem cell gene markers to explore known gaps in annotation transfer methods, illustrating how structure-based deep learning methods such as DeepFRI can significantly improve the functional annotation of heretofore unannotated i-cell markers. Availability and implementation:The Hydractinia Genome Project Portal is freely available at https://research.nhgri.nih.gov/hydractinia.
The epithelial and interstitial stem cells of the freshwater polyp Hydra are the best-characterized stem cell systems in any cnidarian, providing valuable insight into cell type evolution and the origin of stemness in animals. However, little is known about the transcriptional regulatory mechanisms that determine how these stem cells are maintained and how they give rise to their diverse differentiated progeny. To address such questions, a thorough understanding of transcriptional regulation in Hydra is needed. To this end, we generated extensive new resources for characterizing transcriptional regulation in Hydra, including new genome assemblies for Hydra oligactis and the AEP strain of Hydra vulgaris, an updated whole-animal single-cell RNA-seq atlas, and genome-wide maps of chromatin interactions, chromatin accessibility, sequence conservation, and histone modifications. These data revealed the existence of large kilobase-scale chromatin interaction domains in the Hydra genome that contain transcriptionally coregulated genes. We also uncovered the transcriptomic profiles of two previously molecularly uncharacterized cell types: isorhiza-type nematocytes and somatic gonad ectoderm. Finally, we identified novel candidate regulators of cell type-specific transcription, several of which have likely been conserved at least since the divergence of Hydra and the jellyfish Clytia hemisphaerica more than 400 million years ago.
Hydractinia is a colonial marine hydroid that exhibits remarkable biological properties, including the capacity to regenerate its entire body throughout its lifetime, a process made possible by its adult migratory stem cells, known as i-cells. Here, we provide an in-depth characterization of the genomic structure and gene content of two Hydractinia species, H. symbiolongicarpus and H. echinata, placing them in a comparative evolutionary framework with other cnidarian genomes. We also generated and annotated a single-cell transcriptomic atlas for adult male H. symbiolongicarpus and identified cell type markers for all major cell types, including key i-cell markers. Orthology analyses based on the markers revealed that Hydractinia's i-cells are highly enriched in genes that are widely shared amongst animals, a striking finding given that Hydractinia has a higher proportion of phylum-specific genes than any of the other 41 animals in our orthology analysis. These results indicate that Hydractinia's stem cells and early progenitor cells may use a toolkit shared with all animals, making it a promising model organism for future exploration of stem cell biology and regenerative medicine. The genomic and transcriptomic resources for Hydractinia presented here will enable further studies of their regenerative capacity, colonial morphology, and ability to distinguish self from non-self.
Although genomic research has predominantly relied on phenotypic ascertainment of individuals affected with heritable disease, the falling costs of sequencing allow consideration of genomic ascertainment and reverse phenotyping (the ascertainment of individuals with specific genomic variants and subsequent evaluation of physical characteristics). In this research modality, the scientific question is inverted: investigators gather individuals with a genomic variant and test the hypothesis that there is an associated phenotype via targeted phenotypic evaluations. Genomic ascertainment research is thus a model of predictive genomic medicine and genomic screening. Here, we provide our experience implementing this research method. We describe the infrastructure we developed to perform reverse phenotyping studies, including aggregating a super-cohort of sequenced individuals who consented to recontact for genomic ascertainment research. We assessed 13 studies completed at the National Institutes of Health (NIH) that piloted our reverse phenotyping approach. The studies can be broadly categorized as (1) facilitating novel genotype-disease associations, (2) expanding the phenotypic spectra, or (3) demonstrating ex vivo functional mechanisms of disease. We highlight three examples of reverse phenotyping studies in detail and describe how using a targeted reverse phenotyping approach (as opposed to phenotypic ascertainment or clinical informatics approaches) was crucial to the conclusions reached. Finally, we propose a framework and address challenges to building collaborative genomic ascertainment research programs at other institutions. Our goal is for more researchers to take advantage of this approach, which will expand our understanding of the predictive capability of genomic medicine and increase the opportunity to mitigate genomic disease.
Modulation of metabolic flux through pyruvate dehydrogenase complex (PDC) plays an important role in T cell activation and differentiation. PDC sits at the transition between glycolysis and the tricarboxylic acid cycle and is a major producer of acetyl-CoA, marking it as a potential metabolic and epigenetic node To understand the role of pyruvate dehydrogenase complex in T cell differentiation, we generated mice deficient in T cell pyruvate dehydrogenase E1A (Pdha) subunit using a CD4-cre recombinase-based strategy. Herein, we show that genetic ablation of PDC activity in T cells (TPdh-/-) leads to marked perturbations in glycolysis, the tricarboxylic acid cycle, and OXPHOS. TPdh-/- T cells became dependent upon substrate level phosphorylation via glycolysis, secondary to depressed OXPHOS. Due to the block of PDC activity, histone acetylation was also reduced, including H3K27, a critical site for CD8+ TM differentiation. Transcriptional and functional profiling revealed abnormal CD8+ TM differentiation in vitro. Collectively, our data indicate that PDC integrates the metabolome and epigenome in CD8+ memory T cell differentiation. Targeting this metabolic and epigenetic node can have widespread ramifications on cellular function.
Acute myeloid leukemia (AML) is characterized by recurring chromosomal abnormalities that encode oncogenic fusion proteins. Approximately 10% of AML cases are associated with a chromosome 16 inversion [inv(16)(p13q22)] or translocation t(16;16)(p13q22). These chromosomal rearrangements results in the formation of the fusion oncogene CBFB-MYH11 and ultimately the fusion protein CBFβ-SMMHC. Although CBFβ-SMMHC was initially considered a dominant negative repressor of RUNX1, we have recently shown that it works together with RUNX1 to activate gene expression through direct target gene binding. RUNX1 is known to regulate gene expression at both transcriptional and epigenetic levels. Aberrant DNA methylation patterns have been reported across different cancer types including AML. Recent studies have shown that DNA methylation negatively influences RUNX1 DNA binding. We hypothesize that changes to DNA methylation and consequently dysregulated RUNX1 DNA binding contribute to the pathogenesis of inv(16) AML. In this study, we examined the epigenetic and transcriptomic landscapes of CBFB-MYH11 knockin mice and assessed the contribution of the resulting fusion protein CBFβ-SMMHC to RUNX1 DNA target binding. To assess the effects of DNA methylation on RUNX1 DNA binding in vitro, we initially performed fluorescence polarization and electrophoretic mobility shift assays using unmethylated and methylated RUNX1 DNA target sequences. RUNX1, RUNX1/CBFβ and RUNX1/CBFβ-SMMHC showed significantly lower binding affinity for methylated DNA compared to unmethylated DNA probes. To determine genome-wide DNA methylation changes, we performed enzymatic methyl sequencing on sorted Lin -Sca1 -c-Kit + bone marrow cells from control and pre-leukemic inv(16) mice. Inv(16) mice showed a significant genome-wide DNA hypermethylation compared to controls (13,478 hyper- vs 705 hypomethylated regions; q-value < 0.05), with 14% of the differentially methylated regions (DMRs) lying in promoters. Genes associated with these DMRs were enriched for gene ontology (GO) processes related to hematopoiesis, such as myeloid cell homeostasis and erythrocyte differentiation, and DNA replication, such as mitotic DNA replication initiation, among others. In addition, several RUNX1 target genes showed promoter hypermethylation, such as Klf1 or the thrombopoietin receptor Mpl. To understand the impact of aberrant DNA methylation on gene expression, we performed RNA-seq using the same cell population. Gene expression changes were observed in inv(16) mice compared to controls, with 741 genes being upregulated and 278 downregulated from a total of 1019 differentially expressed genes (DEG; absolute fold-change ≥ 2, q-value < 0.05). Enriched GO processes identified from the DEGs were associated with immune responses, such as response to interferon-alpha (and beta) and leukocyte activation involved in inflammatory response. We then intersected the DMR (focusing on promoters) and DEG lists to identify genes whose expression might be directly influenced by DNA methylation changes in inv(16) leukemia. Hypermethylation of promoters was associated with 102 genes, with 67% of them showing downregulation of gene expression in inv(16) mice, whereas hypomethylation was associated with 18 genes with 94% showing upregulation of gene expression. These results show that in inv(16) mice, promoter hypermethylation is associated with downregulation of gene expression, and hypomethylation with increased expression. We also confirmed that several RUNX1 target genes showing promoter hypermethylation also had reduced gene expression in inv(16) mice. The initial analysis suggests that CBFβ-SMMHC expression significantly affects the methylation landscape in mice, with important RUNX1 target genes exhibiting promoter hypermethylation and subsequent downregulation of gene expression. We are currently performing chromatin immunocleavage sequencing using the same cell population to determine RUNX1 binding in vivo. Overall, this study explores a novel regulatory mechanism of RUNX1 function through DNA methylation, and ultimately a new role for RUNX1 in the pathogenesis of inv(16) AML.
The epithelial and interstitial stem cells of the freshwater polyp Hydra are the best characterized stem cell systems in any cnidarian, providing valuable insight into cell type evolution and the origin of stemness in animals. However, little is known about the transcriptional regulatory mechanisms that determine how these stem cells are maintained and how they give rise to their diverse differentiated progeny. To address such questions, a thorough understanding of transcriptional regulation in Hydra is needed. To this end, we generated extensive new resources for characterizing transcriptional regulation in Hydra , including new genome assemblies for Hydra oligactis and the AEP strain of Hydra vulgaris , an updated whole-animal single-cell RNA-seq atlas, and genome-wide maps of chromatin interactions, chromatin accessibility, sequence conservation, and histone modifications. These data revealed the existence of large chromatin interaction domains in the Hydra genome that likely influence transcriptional regulation in a manner distinct from topologically associating domains in bilaterians. We also uncovered the transcriptomic profiles of two previously molecularly uncharacterized cell types, isorhiza-containing nematocytes and somatic gonad ectoderm. We identified novel candidate regulators of cell-type-specific transcription, several of which have likely been conserved at least since the divergence of Hydra and the jellyfish Clytia hemisphaerica over 200 million years ago. The resources generated in this study, which collectively represent the most comprehensive characterization of transcriptional regulation in a cnidarian to date, are accessible through a newly created genome portal, available at research.nhgri.nih.gov/HydraAEP/ .
To address the void in the availability of high-quality proteomic data traversing the animal tree, we have implemented a pipeline for generating de novo assemblies based on publicly available data from the NCBI Sequence Read Archive, yielding a comprehensive collection of proteomes from 100 species spanning 21 animal phyla. We have also created the Animal Proteome Database (AniProtDB), a resource providing open access to this collection of high-quality metazoan proteomes, along with information on predicted proteins and protein domains for each taxonomic classification and the ability to perform sequence similarity searches against all proteomes generated using this pipeline. This solution vastly increases the utility of these data by removing the barrier to access for research groups who do not have the expertise or resources to generate these data themselves and enables the use of data from nontraditional research organisms that have the potential to address key questions in biomedicine.
How individuals perceive uncertainties in sequencing results may affect their clinical utility. The purpose of this study was to explore perceptions of uncertainties in carrier results and how they relate to psychological well-being and health behavior. Post-reproductive adults (N = 462) were randomized to receive carrier results from sequencing through either a web platform or a genetic counselor. On average, participants received two results. Group differences in affective, evaluative, and clinical uncertainties were assessed from baseline to 1 and 6 months; associations with test-specific distress and communication of results were assessed at 6 months. Reductions in affective uncertainty (∆x̅ = 0.78, 95% CI: 0.53, 1.02) and evaluative uncertainty (∆x̅ = 0.69, 95% CI: 0.51, 0.87) followed receipt of results regardless of randomization arm at 1 month. Participants in the web platform arm reported greater clinical uncertainty than those in the genetic counselor arm at 1 and 6 months; this was corroborated by the 1,230 questions asked of the genetic counselor and residual questions reported by those randomized to the web platform. Evaluative uncertainty was associated with a lower likelihood of communicating results to health care providers. Clinical uncertainty was associated with a lower likelihood of communicating results to children. Learning one’s carrier results may reduce perceptions of uncertainties, though web-based return may lead to less reduction in clinical uncertainty in the short term. These findings warrant reinforcement of clinical implications to minimize residual questions and promote appropriate health behavior (communicating results to at-risk relatives in the case of carrier results), especially when testing alternative delivery models.
Adeno-associated viral (AAV) vectors have emerged as the preferred platform for in vivo gene transfer because of their combined efficacy and safety. However, insertional mutagenesis with the subsequent development of hepatocellular carcinomas (HCCs) has been recurrently noted in newborn mice treated with high doses of AAV, and more recently, the association of wild-type AAV integrations in a subset of human HCCs has been documented. Here, we address, in a comprehensive, prospective study, the long-term risk of tumorigenicity in young adult mice following delivery of single-stranded AAVs targeting liver. HCC incidence in mice treated with therapeutic and reporter AAVs was low, in contrast to what has been previously documented in mice treated as newborns with higher doses of AAV. Specifically, HCCs developed in 6 out 76 of AAV-treated mice, and a pathogenic integration of AAV was found in only one tumor. Also, no evidence of liver tumorigenesis was found in juvenile AAV-treated mucopolysaccharidosis type VI (MPS VI) cats followed as long as 8 years after vector administration. Together, our results support the low risk of tumorigenesis associated with AAV-mediated gene transfer targeting juvenile/young adult livers, although constant monitoring of subjects enrolled in AAV clinical trial is advisable.
For over a thousand years, the common goldfish (Carassius auratus) was raised throughout Asia for food and as an ornamental pet. As a very close relative of the common carp (Cyprinus carpio), goldfish share the recent genome duplication that occurred approximately 14 million years ago in their common ancestor. The combination of centuries of breeding and a wide array of interesting body morphologies provides an exciting opportunity to link genotype to phenotype and to understand the dynamics of genome evolution and speciation. We generated a high-quality draft sequence and gene annotations of a "Wakin" goldfish using 71X PacBio long reads. The two subgenomes in goldfish retained extensive synteny and collinearity between goldfish and zebrafish. However, genes were lost quickly after the carp whole-genome duplication, and the expression of 30% of the retained duplicated gene diverged substantially across seven tissues sampled. Loss of sequence identity and/or exons determined the divergence of the expression levels across all tissues, while loss of conserved noncoding elements determined expression variance between different tissues. This assembly provides an important resource for comparative genomics and understanding the causes of goldfish variants.
Single-cell RNA sequencing (scRNA-seq) of human primary tissues is a rapidly emerging tool for investigating human health and disease at the molecular level. However, optimal processing of solid tissues presents a number of technical and logistical challenges, especially for tissues that are only available at autopsy, which includes pancreatic islets, a tissue that is highly relevant to diabetes. To assess the possible effects of different sample preparation protocols on fresh islet samples, we performed a detailed comparison of scRNA-seq data generated with islets isolated from a human donor but processed according to four treatment strategies, including fixation and cryopreservation. We found significant and reproducible differences in the proportion of cell types identified, and more minor effects on cell-specific patterns of gene expression. Fresh islets from a second donor confirmed gene expression signatures of alpha and beta subclusters. These findings may well apply to other tissues, emphasizing the need for careful consideration when choosing processing methods, comparing results between different studies, and/or interpreting data in the context of multiple cell types from preserved tissue.
The ADAM (a disintegrin and metalloprotease) protein family uniquely exhibits both catalytic and adhesive properties. In the well-defined process of ectodomain shedding, ADAMs transform latent, cell-bound substrates into soluble, biologically active derivatives to regulate a spectrum of normal and pathological processes. In contrast, the integrin ligand properties of ADAMs are not fully understood. Emerging models posit that ADAM–integrin interactions regulate shedding activity by localizing or sequestering the ADAM sheddase. Interestingly, 8 of the 21 human ADAMs are predicted to be catalytically inactive. Unlike their catalytically active counterparts, integrin recognition of these “dead” enzymes has not been largely reported. The present study delineates the integrin ligand properties of a group of non-catalytic ADAMs. Here we report that human ADAM11, ADAM23, and ADAM29 selectively support integrin α4-dependent cell adhesion. This is the first demonstration that the disintegrin-like domains of multiple catalytically inactive ADAMs are ligands for a select subset of integrin receptors that also recognize catalytically active ADAMs.
It is unclear how standing genetic variation affects the prognosis of prostate cancer patients. To provide one controlled answer to this problem, we crossed a dominant, penetrant mouse model of prostate cancer to Diversity Outbred mice, a collection of animals that carries over 40 million SNPs. Integration of disease phenotype and SNP variation data in 493 F1 males identified a metastasis modifier locus on Chromosome 8 (LOD = 8.42); further analysis identified the genes Rwdd4, Cenpu, and Casp3 as functional effectors of this locus. Accordingly, analysis of over 5,300 prostate cancer patient samples revealed correlations between the presence of genetic variants at these loci, their expression levels, cancer aggressiveness, and patient survival. We also observed that ectopic overexpression of RWDD4 and CENPU increased the aggressiveness of two human prostate cancer cell lines. In aggregate, our approach demonstrates how well-characterized genetic variation in mice can be harnessed in conjunction with systems genetics approaches to identify and characterize germline modifiers of human disease processes.
New sequencing technologies are becoming increasingly available in a variety of contexts. These techniques include whole-exome (we refer to this as “exome” sequencing) and whole-genome sequencing (Biesecker and Green 2014). These types of genomic sequencing have impacted both research and clinical practice. Genomic sequencing has led to the discovery of novel genetic etiologies of disease and has also improved the ability to diagnose patients with subtle or atypical presentations of genetic conditions (especially those affected by relatively rare disorders) by allowing simultaneous interrogation of many loci (Boycott et al. 2013; Yang et al. 2013, 2014; Taylor et al. 2015). To quantify genetic progress and the impact of these technologies in understanding the causes of Mendelian disorders, we analyzed methods of discovery in the last ~2.5 years. We reviewed all newly described Mendelian disease genes in the ~2.5-year period (April 30, 2013–November 30, 2015) following the initial public dissemination of the Clinical Genomic Database (CGD) (Solomon et al. 2013), a freely available web-based resource that focuses on the clinical sequelae and management of genetic disorders (the CGD, which is regularly updated to keep pace with genetic knowledge, is available at: http://research.nhgri.nih.gov/CGD/). Regarding relevant literature, as there is frequently a gap between manuscript acceptance, electronic, and final publication, we attempted to include only those genes/conditions in the ~2.5 year period that were published (available in PubMed) and subsequently included in both the CGD and Online Mendelian Inheritance in Man (OMIM, available at http://www.omim.org) within that ~2.5 year interval. We recognize this is imperfect for several reasons. These reasons include the lag between publication and incorporation in these databases, and the fact that the databases do not completely capture the literature and are themselves being continually being refined and updated. However, with these caveats in terms of potential inaccuracies, the trends that we reveal are overall interesting. For each article, we determined how the discovery was made (e.g., through whole-genome sequencing alone, homozygosity mapping and exome sequencing, candidate gene studies, or a combination of possibilities). If the discovery involved findings through a previous publication on the same families or condition (such as linkage analysis implicating specific loci), we included that previous method as part of the discovery process. In the example given, this would be considered “exome + linkage.” We also determined if the discovery was made through analysis of a single individual; a single family; or a single individual or family followed by the identification of additional mutation-positive individuals through studies of a larger cohort. Finally, we investigated whether additional bench-based studies, aimed at providing understanding and evidence beyond clinical and bioinformatics results, were included in the investigation. We did not include newly reported conditions allelic to previously described conditions. Regarding this latter exclusion criterion, though many genes have clinically distinct allelic disorders, we wished to be maximally conservative in order to avoid any controversy arising from similar disorders that may represent the spectrum of a single disease entity. For newly discovered genetic causes, 445 new genes were identified in the 2.5-year period in the defined intervals (see Fig. 1 and Table S1 for details). A number of genes were identified by apparently independent studies with separate but simultaneous publications. Methods from each such independent study were counted separately here, such that 492 individual studies are represented. Of the 492 studies, 394 (80%) used some type of genomic sequencing (including exome or genome sequencing but not X-chromosome exome or mitochondrial exome). Three hundred and seventy-nine (77%) used exome sequencing, 11 (2%) used genome sequencing, and four (1%) used both exome and genome sequencing in the same study. Two hundred and thirty-one (47%) used only exome sequencing (without other investigations such as homozygosity mapping) for gene identification; 238 (48%) used only exome and/or genome sequencing. Another 156 (32%) used exome or genome sequencing in addition to other techniques, such as homozygosity mapping. Forty-one (8%) used a traditional candidate gene approach without other methods, though 46 (10%) used a candidate gene approach with methods other than exome or genome sequencing. The remaining 11 (2%) used other methods, such as cytogenomic methods alone, X-chromosome exome sequencing, or mitochondrial sequencing (see Table S1). Eighteen (4%) of the overall studies investigated a single patient only; 114 (23%) studied multiple members of a single family; 93 (19%) were initially done on a single patient or family but then investigated a larger cohort based on these initial findings and identified and reported additional mutation-positive individuals. The remaining 267 (54%) were reported as studying multiple patients or families simultaneously. We would predict that over time, the use of genomic sequencing methods will eclipse other methods of identifying disease genes. To detect if there has been a shift in methods, even over the short ~2.5 year period of this study, we sorted the 492 publications by PubMed identification number (PMID) and split the publications into two groups. We recognize that the PMID identifiers do not precisely capture chronology. Nevertheless, we note that studies in the second group, which represents manuscripts published more recently, made significantly more use of genomic sequencing overall (210 vs. 184 studies, P = 0.0046 by Fisher's exact test) and genomic sequencing alone (135 vs. 103 studies, P = 0.0051 by Fisher's exact test) than did those in the first. Of the 492 studies, 359 (73%) included some type of laboratory-based assay in addition to standard clinical work-up and genomic sequencing and analysis. However, we intentionally did not further analyze the type of cellular or functional analyses that were done as part of the paper – we viewed this assignment as difficult and potentially unhelpful to interpret, as this depended on the availability of previous knowledge about a particular gene's function (e.g., through an animal model), the specific mutation in question (e.g., a truncating vs. missense mutation), may have been aimed at understanding the biological implications of the genetic pathways involved rather than or in addition to the question of mutation pathogenicity, and because so many different possible methods might have been used (MacArthur et al. 2014). Our analyses show that exome sequencing currently accounts for the vast majority of causal gene discovery. Almost half of all the discoveries investigated here occurred through exome sequencing alone. This is anticipated to shift to genome sequencing as affordability and accuracy of sequencing continue to improve (Hayden 2014). Additionally, our analyses show a statistically significant increase in the use of exome or genome sequencing even within the short time period analyzed. Again, though we admit to potential inaccuracies as the background databases shift, we feel that the trends described here are illustrative. In conclusion, new sequencing technologies are resulting in a dramatic change in the discovery of the causes of disease. These discoveries can be quickly translated into clinical care – patients affected with many conditions now have a better chance of explanations based on genetic testing. Further, finding the molecular etiology of a condition may then in turn yield more tailored and overall better management (Solomon et al. 2013; Soden et al. 2014; Khromykh and Solomon 2015; Solomon 2015; Willig et al. 2015). It will be interesting to perform similar analyses in the future as new techniques emerge and as the body of knowledge of the causes of human disease continues to grow and evolve. This research was supported in part by the Intramural Research Program of the National Human Genome Research Institute, the National Institutes of Health. None of the authors have disclosures or conflicts of interest. Please note: The publisher is not responsible for the content or functionality of any supporting information supplied by the authors. Any queries (other than missing content) should be directed to the corresponding author for the article.