Accurate estimates of allele frequencies aid in genetic discovery, including rare disease diagnosis, common disease investigations, and population genetics. Here, we present the Genome Aggregation Database version 4 (gnomAD v4), including 730,947 with exome sequences, a fivefold increase over previous releases. We demonstrate that statistical power to detect strong selective constraint continues to increase with sample size. We develop a new loss-of-function annotation pipeline, which learns genomic features predictive of nonsense-mediated decay and splicing effects from selection signals, achieving 90% precision for distinguishing likely true versus false positive loss-of-function variants. This improved pipeline, along with incorporation of highly deleterious missense variants into measures of loss-of-function intolerance, improves disease gene detection, particularly for short genes and those with gain-of-function mechanisms. To improve disease gene prediction, we systematically extract gene-disease associations from biomedical literature, map these to gene-level biological features, and integrate both with refined constraint metrics within a Bayesian framework, yielding state-of-the-art prediction of gene-disease relevance. We highlight genes under strong constraint but with limited clinical characterization, which are enriched in embryonic lethal and fertility phenotypes, thus prioritizing previously under-characterized disease genes. Together, these advances establish a unified framework for accelerating gene discovery and improving rare disease diagnosis.
One of the seminal discoveries from genetic studies of autism spectrum disorder and related neurodevelopmental disorders (NDDs) has been that loss-of-function (LoF) mutations in genes that impact transcriptional regulation confer substantial liability to NDDs. Haploinsufficiency of the epigenetic regulator POGZ represents one of the strongest such associations; however, little is known about the mechanisms by which POGZ LoF alters early neuronal development. Here, we created an allelic series of CRISPR-engineered human induced pluripotent stem cell (hiPSC) clones harboring mono- and bi-allelic POGZ deletions. In hiPSC-derived neural stem cells (NSCs) and Neurogenin-2-induced neurons (iNs), POGZ LoF altered the expression of genes associated with synaptic and intracellular signaling and extracellular matrix organization. Our multiomics profiling also showed altered footprinting of critical transcription factors (e.g., activator protein 1 complexes) that were enriched at promoters of differentially expressed genes associated with synaptic function. To further interrogate the shared molecular changes associated with NDDs, we compared our results to deletions of the transcription factor MEF2C and the sodium channel gene SCN2A that we generated in these same isogenic iNs. These analyses revealed strong enrichment of extracellular matrix and intracellular signaling disruption associated with POGZ and MEF2C deletion, whereas POGZ and SCN2A haploinsufficiency exhibited shared transcriptional effects on gene modules enriched for NDD-associated genes with opposing regulatory effects. Notably, we also observed alterations to synaptic firing rate and neurite extension with bi-allelic deletions. These shared molecular consequences suggest key points of convergence that connect gene regulation to neuronal function in the etiology of neurodevelopmental pathologies.
The human cortex acquires its advanced cognitive capacity through tightly regulated developmental programs, disruption of which underlies neurodevelopmental disorders such as Schaaf-Yang syndrome (SYS) and Prader-Willi syndrome (PWS). While SYS results from pathogenic variants in the imprinted gene MAGEL2, PWS arises from chromosomal deletions, imprinting defects or uniparental disomy encompassing the MAGEL2 locus. However, the contribution of MAGEL2 to disease pathogenesis and human corticogenesis is not fully understood. Here, we performed integrated transcriptomic, proteomic, and ubiquitinomic profiling of cortical neurons derived from CRISPR/Cas9-engineered isogenic human pluripotent stem cells (hiPSC) modeling SYS and PWS. Beyond PWS-specific signatures including dysregulated ribosomal processes, we identified MAGEL2-dependent defects shared across both disorders. These include reduced progenitor proliferation, accelerated neuronal maturation, impaired migration and adhesion, as well as abnormal synaptic development, collectively linking PWS and SYS at the level of cortical development. Notably, these phenotypes partially overlap with those observed in other neurodevelopmental disorders, suggesting that MAGEL2 governs core pathways broadly vulnerable in disease. Together, our findings establish MAGEL2 as a key regulator of human cortical development, provide a unifying mechanistic framework for SYS and PWS, accessible via a web-based platform.
The NeuroDev study, conducted in Kenya and South Africa, is a large-scale clinical, genetic, and epidemiologic characterization of neurodevelopmental disorders (NDDs) on the African continent. NeuroDev assessments capture birth, demographic, and developmental history; cognitive and behavioral outcomes; and physical health variables. DNA samples are collected for exome sequencing and clinical genetic analysis. This paper presents novel data from 521 children with NDDs, 739 of those children's parents, and 255 unrelated, typically-developing children. The analyses offer unique genetic and phenotypic characterizations of NDDs in two African countries and underscore the importance of including underrepresented populations in NDD research. Ultimately, 107 children with NDDs from the NeuroDev cohort (22.1%) had likely pathogenic or pathogenic variants in established NDD genes. High rates of genetic diagnosis were associated with high rates of environmental risk factors for NDDs. All data, materials, and measures generated from this study are publicly available through the US National Institute of Mental Health.
Rare disease research and diagnosis rely on the integration of genomic and phenotypic data generated across diverse clinical sites; however, the absence of widely adopted standards for representing genomic data and associated metadata has limited data interoperability, reuse, and cross-study analysis. The Genomics Research to Elucidate the Genetics of Rare Diseases (GREGoR) Consortium was established to investigate challenging rare disease cases and evaluate emerging multi-omic technologies for clinical translation. To support coordinated data integration across distributed research sites, we developed a common Consortium Data Model in partnership with domain experts to standardize the capture of participant-, family-, phenotype- and assay-level metadata, with a particular emphasis on using a modular architecture to support linking of multiple data versions from multiple omic technologies to a single individual and attribution of a genetic finding to the specific technology used for its initial discovery. Adoption of the GREGoR Data Model has enabled continued generation and public release of a harmonized, analysis-ready Consortium Dataset. The most recent release includes phenotypic, family and multi-omic data from 12,292 participants in 5,029 families. Other rare disease data sharing efforts are beginning to adopt this data model which will facilitate cross consortium analyses and empower rare disease research. This work demonstrates that a collaborative, flexible, and scalable data model can enable large-scale rare disease research, facilitate cross-center data harmonization, and enable data interoperability.
Cohesin orchestrates gene expression via three-dimensional chromosome folding. Genes encoding cohesin and cohesin loaders have been associated with Mendelian disorders, whereas genes encoding cohesin release factors, including WAPL and its binding partners PDS5A and PDS5B, have not. We explored the relevance of cohesin release factors in Mendelian disease by phenotyping individuals with heterozygous predicted damaging variants in WAPL (n = 27), PDS5A (n = 8), and PDS5B (n = 8), by modeling WAPL deficiency in human cells and mice, and by aggregating disease association statistics from consortia studies. We identified a WAPL-related disorder featuring developmental delay, intellectual disability, and risk of other developmental anomalies. Similarities between individuals with damaging WAPL variants and those with large, recurrent 10q22.3q23.2 (10q) deletions encompassing WAPL nominate WAPL as a driver gene within this genomic disorder region. While individuals with PDS5A or PDS5B variants exhibited features of developmental disorders, neither cohort-based statistics nor subject phenotyping associated these genes with specific phenotypes. We used CRISPR to generate truncating variants in WAPL and 10q deletion or duplication in human induced pluripotent stem cells (iPSCs) and induced neurons. Transcriptomics identified significant overlap between WAPL haploinsufficiency and 10q deletion differentially expressed genes. Mice with 50% Wapl expression exhibited mild deficits of growth and learning/memory, whereas those with 25% residual Wapl displayed birth defects and postnatal lethality, revealing a dosage liability threshold below the level of heterozygosity. In summary, we delineated a genetic condition caused by cohesin release factor deficiency, nominated WAPL as a driver gene within a genomic disorder region, and further illuminated dosage sensitivity of human cohesin.
Rare diseases collectively affect 1 in 10 individuals, yet current genetic testing fails to identify a causal variant for most cases. At present, cytogenetic methods and/or sequencing approaches such as exome (ES) or short-read genome sequencing (srGS) represent the state-of-the-art for comprehensive clinical discovery of sequence and structural variants (SVs), including copy number variants, balanced SVs, complex SVs, and tandem repeats (TRs). Recently, long-read genome sequencing (lrGS), coupled with multiomics data, has presented great promise to resolve variation in genomic regions recalcitrant to characterization by srGS such as highly repetitive simple repeat sequences and segmental duplications. However, there are few guidelines to enable clinical interpretation of genetic variation in these highly repetitive genomic regions, and the enthusiasm of the field in adopting lrGS has made it difficult to assess the true added diagnostic yield of this technology due to widely variable and inconsistently applied analytic pipelines and variable degrees of pre-screening by ES or srGS. Here, we investigated the contribution of SVs to rare diseases using srGS as a front-line strategy when paired with highly sensitive SV discovery and evaluate the added diagnostic yield of incorporating lrGS for a subset of cases. Our srGS analysis encompassed 1,462 families (3,450 individuals) recruited through the Broad Institute Center for Mendelian Genetics and the Genomics Research to Elucidate the Genetics of Rare Diseases (GREGoR) programs. Diagnostic SVs were identified in 5.4% of cases (79/1,462), of which 80% were uniquely detectable by srGS compared to standard cytogenetic techniques. For 96 families (including 10 families with a heterozygous variant observed in a known recessive gene of clinical relevance), we performed lrGS with methylation profiling, as well as long-read transcriptomic analyses in a subset of 20 trios. Analyses with lrGS yielded over 25,000 SVs per genome, 63% of which were not captured by srGS, along with an additional ~200 rare SNV/indels per genome not previously captured and 12 differentially methylated regions per genome. Among these, we identified only one diagnostic variant not interpreted by srGS, an apparently mosaic de novo SNV in CASK that was absent in the srGS callset due to allelic imbalance. No new diagnoses were supported by long-read transcriptomics or episignatures. In this well characterized rare disease cohort, the added diagnostic yield was thus 1.04% (1/96 families). Following a systematic literature review of prior lrGS studies, we find that most reported diagnoses were detectable by srGS and that our added diagnostic yield is consistent with those prior studies. These studies emphasize the significant impact of comprehensive SV discovery in rare disease cases and further demonstrate the power for increased discovery of novel genomic variation and episignatures from lrGS. Nonetheless, they also serve to temper expectations of dramatic diagnostic advances in rare disease patients until there is more extensive annotation of the functional and clinical impact of all coding and noncoding variation uniquely accessible to lrGS with extensive reference databases spanning highly repetitive genomic sequencing that could be enabled by this transformative technology.
Cohesin is a fundamental genome-organizing complex that orchestrates three-dimensional chromosome folding and gene expression via DNA loop extrusion. Alterations to genes encoding cohesin subunits and cohesin loaders cause Mendelian disorders, including Cornelia de Lange syndrome (CdLS). By contrast, disruption of factors that remove cohesin from DNA, including WAPL and its binding partners PDS5A and PDS5B, have not yet been associated with human disease. Here, we explored the relevance of these cohesin release factors in Mendelian disease by establishing a rare disease cohort of deeply phenotyped individuals with heterozygous, predicted damaging variants in WAPL (n=27), PDS5A (n=8), and PDS5B (n=8), by modeling WAPL deficiency in human cell lines and mice, and by aggregating rare disease association statistics from consortia studies. We identified a WAPL-related disorder characterized by developmental delay, intellectual disability, and risk of other developmental anomalies including clubfoot. Similarities between individuals with damaging WAPL variants and those with large, recurrent 10q22.3q23.2 (10q) deletions (which encompass WAPL) nominate WAPL as a driver gene within this genomic disorder region. While carriers of PDS5A or PDS5B variants exhibited features of developmental disorders, neither cohort-based statistics nor case phenotyping associated these genes with specific phenotypes. We used CRISPR engineering to generate truncating variants in WAPL, as well the 7.8 Mb 10q deletion or duplication in human iPSCs and induced neurons. Transcriptomic analyses identified differentially expressed genes in both models, with highly significant overlap between WAPL haploinsufficiency and 10q deletion signatures. Mice with 50% residual Wapl expression exhibited mild deficits of growth and learning/memory, whereas those with 25% residual Wapl expression displayed birth defects and postnatal lethality, revealing a dosage liability threshold below the level of heterozygosity. In summary, we delineated a novel genetic condition caused by cohesin release factor deficiency, nominated WAPL as a driver gene within a genomic disorder region, and further illuminated dosage sensitivity of human cohesin.
Here we developed and deployed the blended genome exome (BGE) method, a DNA library approach that generates low-pass whole-genome (1-4× mean depth) and deep whole-exome (30-40× mean depth) data in a single sequencing run. BGE is cost-effective, empowers most genomic discoveries possible with deep whole-genome sequencing and captures global common single-nucleotide polymorphism diversity. We applied BGE to sequence >53,000 samples from the PUMAS Project (Populations Underrepresented in Mental Illness Associations Studies), including African, African American and Latin American populations. Imputed genotypes showed high concordance with Illumina Global Screening Array calls (R2 ≥ 95% for minor allele frequency ≥1%; ≥90% for minor allele frequency <1%), with consistent performance across local ancestries in admixed cohorts. For protein-coding copy number variants, deletions and duplications spanning at least three exons had a positive predicted value of ~90% relative to deep whole-genome data. At ~28% of the cost of deep whole-genome sequencing, BGE provides a scalable, reliable platform to expand genomic discovery and equitable access to sequencing in underrepresented populations.
Genome-wide association studies (GWAS) and large-scale rare variant burden analyses have identified both common and rare loss-of-function variants associated with neuropsychiatric and neurodegenerative disorders. Yet, the shared biological processes influenced by both classes of variation remain poorly characterized. In this study, we utilized transcriptomic data from 933 post-mortem brain samples to identify genes that show convergent coexpression with GWAS and rare variant burden risk genes across six brain disorders. Despite largely distinct sets of significant risk genes from GWAS and rare variant burden studies, we found a significant overlap in their convergently coexpressed genes. These convergent genes showed enrichment for common and rare variant heritability and highlighted key biological pathways and cell-type markers impacted by both types of genetic variation. Compared to genes coexpressed with one variant class, shared convergent genes exhibited stronger evolutionary constraint and greater enrichment for known drug targets, underscoring their potential therapeutic relevance. Collectively, our results establish a systematic and generalizable framework for integrating coexpression data with genetic risk to reveal transcriptional programs supported by both common and rare variant evidence, offering mechanistic insights into neuropsychiatric diseases.
Autism spectrum disorder is a heritable neurodevelopmental condition affecting approximately 3% of children that presents with core behavioral features and a range of possible comorbidities, including intellectual disability. While common variants contribute substantially to autism liability, the discovery of specific autism-associated genes has largely been driven by studies of rare and de novo variants. Many of these genes are also linked with broadly defined developmental disorders, but their involvement in other conditions has not been mapped at scale. Here, we analyze autosomal rare coding variation from 62,429 individuals with autism from research and clinical cohorts to identify 253 autism-associated genes at an estimated false discovery rate < 0.001. We cluster them based on association evidence from large-scale studies of developmental disorders, schizophrenia, bipolar disorder, and epilepsy, generating six clusters of genes with differing biological pathway enrichments and patterns of comorbidities. Investigating rare variant associations in the population using the UK Biobank and All of Us, we identify autism-associated genes displaying pleiotropy across physiological systems. In addition, we report 497 genes impacting development in a meta-analysis with 26,109 published developmental disorders samples. Collectively drawing upon data from over 1.5 million individuals, our study finds that rare variants across hundreds of genes contribute to autism with variable phenotypic outcomes.
Advances in transcriptomics have transformed our understanding of amyotrophic lateral sclerosis (ALS), a progressive neurodegenerative disease, revealing disrupted gene expression profiles and highlighting the multi-system biology of ALS. Despite major advances, transcriptomic studies have only begun to capture the complexity and the molecular hierarchy of transcriptomic alterations in ALS. To resolve and characterize the transcriptome in ALS, we performed a comprehensive reanalysis of bulk RNA sequencing from the New York Genome Center ALS Consortium cohort across five post-mortem tissues including motor and frontal cortex, cervical and lumbar spinal cord, and cerebellum. By deploying dual analytical pipelines - one reference-based to model canonical events and one de novo to detect transcript structural novelties - we disentangled the quantitative and qualitative architectures of ALS. Our reference-based analysis revealed that ALS transcriptome is defined primarily by splicing failure rather than changes in gene expression. Aberrant splicing events, particularly intron retention, outnumbered differentially expressed genes by an order of magnitude. This widespread loss of fidelity disproportionately affected RNA-binding proteins, suggesting a collapse in their autoregulatory feedback loops. Deconvolution of these signals identified distinct cellular vulnerabilities: transcriptional disruptions were enriched in glial cells in sporadic cases but in neuronal cells in C9ORF72-positive cases. Furthermore, we observed sex-specific dysregulation, with male patients exhibiting greater disruption in guanosine triphosphatase signaling and ciliary organization pathways. In parallel, our de novo analysis uncovered a significant burden of disease-specific gene fusions that were absent in controls. Whole-genome sequencing of the same individuals, together with a larger reference population confirmed that disease-specific fusions do not arise from genomic structural variants, indicating a transcriptional rather than genomic origin. Investigation into the mechanism of these RNA-based fusions revealed a critical deviation in splice site definition: while canonical splice junctions exhibit a high density of binding motifs for polyA-binding or 3'-cleaveage proteins approximately 50 base pairs upstream of the splice donor site (left junction), ALS-specific fusion junctions displayed a dramatic depletion of these motifs in the same region. Functionally, the presence of these sparse disease-specific fusions was strongly correlated with severe splicing outliers in genes governing guanosine triphosphatase activity, converging with the tissue- and male-specific defects identified in our reference-based analysis. Altogether, our results delineated a transcriptome characterized by aberrant splicing with tissue-and sex-specific changes and identified structural-variant-independent RNA fusions as candidate disease modifiers that may amplify pathology. This integrated view provides a mechanistic scaffold for splicing-centered and RNA-structural therapeutic strategies for ALS.
Tauopathies encompass diverse neurodegenerative diseases unified by aberrant patterns of tau deposition in brain. Although most appear sporadic, some are linked to genetic etiologies that offer unique mechanistic insights. Here we report that X-linked Dystonia-Parkinsonism (XDP), caused by a non-coding retrotransposon-associated repeat insertion in TAF1 , involves a significant imbalance of tau isoforms and the accumulation of hyperphosphorylated, four-repeat tau in the brain. In striatal tissue, both misfolded tau accumulation, predominantly in astrocytes, and MAPT exon 10 inclusion correlated with repeat length within the causal insertion. Transcriptomic profiling across brain regions revealed dysregulation of known tau-related pathways. Levels of phosphorylated tau181, glial fibrillary acidic protein, and neurofilament light chain were elevated in patient plasma and discriminated XDP from controls. These findings implicate defective tau proteostasis as a key pathogenic mechanism and position XDP as a genetic model for uncovering cellular drivers that may disrupt tau in other more common neurodegenerative diseases.
This article is based on the address given by the author at the 2025 meeting of The American Society of Human Genetics (ASHG) in Boston, MA. A video of the original address can be found at the ASHG website.
The past decade has seen remarkable progress in identifying genes that, when impacted by deleterious coding variation, confer high likelihood for autism spectrum disorder (ASD), intellectual disability and other associated developmental disorders. However, most underlying gene discovery efforts have focused on individuals of European ancestry, limiting insights into genetic liability across diverse populations. To help address this, the Genomics of Autism in Latin American Ancestries (GALA) Consortium was formed, presenting here the largest sequencing study of autism in Latin American individuals (n > 15,000, including 4,717 participants with an ASD diagnosis). We identified 35 genome-wide significant (false discovery rate < 0.05) autism-associated genes, with substantial overlap with findings from European cohorts, and highly constrained genes showing consistent signal across populations. The results provide support for emerging (for example, MARK2, YWHAG, PACS1, RERE, SPEN, GSE1, GLS, TNPO3 and ANKRD17) and established autism genes and for the utility of genetic testing approaches for deleterious variants in individuals from diverse backgrounds; the results also demonstrate the ongoing need for more inclusive genetic research and testing. We conclude that the biology of autism is consistent across populations, with no detectable influence of ancestry.
Cytogenetic technologies such as G-banding chromosome and FISH analyses have long been the gold standard diagnostic test in prenatal genetic testing. However, unbiased next-generation sequencing technologies such as fetal exome or genome sequencing (ES/GS) are becoming widely accessible and increasingly utilized, particularly for fetuses with structural anomalies. Emerging studies are now establishing increased diagnostic yields from molecular technologies, but there remains a lack of consensus as to whether ES/GS should replace cytogenetic technologies and targeted genepanel screening as first-line tests for all prenatal diagnoses. This report is a summary of the debate on this topic presented at the 28th International Conference on Prenatal Diagnosis and Fetal Therapy. Both expert debaters discussed the advantages and disadvantages.
Postoperative delirium is a type of acute cognitive dysfunction characterized by inattention, disorganized thinking, and altered levels of consciousness that commonly develops after major surgery. Efforts to reduce the incidence of delirium have focused primarily on optimizing perioperative care, however the development of prophylactic interventions have been hindered by a limited understanding of the underlying mechanisms involved in delirium. In this secondary analysis of the Minimizing ICU Neurological Dysfunction with Dexmedetomidine-induced Sleep (MINDDS) trial, a nested case-control study (n = 51) was conducted using total RNA-sequencing analysis of whole-blood to investigate genes associated with delirium risk and development. Transcriptomic analysis revealed significantly lower expression of a key complement pathway inhibitor, C4BPA, in participants who experienced postoperative delirium. This finding was confirmed by quantitative PCR in the MINDDS cohort (n = 319) in adjusted logistic models. Furthermore, complement inhibitor CD55 was also found to be under-expressed in participants who developed delirium. Dexmedetomidine treatment modified associations between C4BPA and CD55 expression and the incidence of postoperative delirium by decreasing incidence in participants with low C4BPA and CD55 expression. This study revealed key complement regulators as risk biomarkers of postoperative delirium. Importantly, our findings suggest postoperative delirium risk is modifiable. Unlike previous research that has mainly focused on proteomics, this study underscores the effectiveness of whole-blood transcriptomics in identifying biomarkers and underlying biological mechanisms of postoperative delirium.
CRISPR-based gene activation (CRISPRa) has emerged as a promising therapeutic approach for neurodevelopmental disorders (NDD) caused by haploinsufficiency. However, scaling this cis -regulatory therapy (CRT) paradigm requires pinpointing which candidate cis -regulatory elements (cCREs) are active in human neurons, and which can be targeted with CRISPRa to yield specific and therapeutic levels of target gene upregulation. Here, we combine Massively Parallel Reporter Assays (MPRAs) and a multiplex single cell CRISPRa screen to discover functional human neural enhancers whose CRISPRa targeting yields specific upregulation of NDD risk genes. First, we tested 5,425 candidate neuronal enhancers with MPRA, identifying 2,422 that are active in human neurons. Selected cCREs also displayed specific, autonomous in vivo activity in the developing mouse central nervous system. Next, we applied multiplex single-cell CRISPRa screening with 15,643 gRNAs to test all MPRA-prioritized cCREs and 761 promoters of NDD genes in their endogenous genomic contexts. We identified hundreds of promoter- and enhancer-targeting CRISPRa gRNAs that upregulated 200 of the 337 NDD genes in human neurons, including 91 novel enhancer-gene pairs. Finally, we confirmed that several of the CRISPRa gRNAs identified here demonstrated selective and therapeutically relevant upregulation of SCN2A , CHD8 , CTNND2 and TCF4 when delivered virally to patient cell lines, human cerebral organoids, and a humanized mouse model of hTcf4 . Our results provide a comprehensive resource of active, target-linked human neural enhancers for NDD genes and corresponding gRNA reagents for CRT development. More broadly, this work advances understanding of neural gene regulation and establishes a generalizable strategy for discovering CRT gRNA candidates across cell types and haploinsufficient disorders.
Rare diseases are collectively common, affecting approximately 1 in 20 individuals worldwide. In recent years, rapid progress has been made in rare disease diagnostics due to advances in next-generation sequencing, development of new computational and functional genomics approaches to prioritize genes and variants and increased global sharing of clinical and genetic data. However, more than half of individuals suspected to have a rare disease lack a genetic diagnosis. The Genomics Research to Elucidate the Genetics of Rare Diseases (GREGoR) Consortium was initiated to study thousands of challenging rare disease cases and families and apply, standardize and evaluate emerging genomics technologies and analytics to accelerate their adoption in clinical practice. Furthermore, all data generated, currently representing over 7,500 individuals from over 3,000 families, are rapidly made available to researchers worldwide through the Analysis, Visualization and Informatics Lab-space (AnVIL) to catalyse global efforts to develop approaches for genetic diagnoses in rare diseases. Most of these families have undergone previous clinical genetic testing but remained unsolved, with most being exome-negative. Here we describe the collaborative research framework, datasets and discoveries comprising GREGoR that will provide foundational resources and substrates for the future of rare disease genomics.
The All of Us Research Program (AoU) is a national biobank seeking to enroll one million individuals in the United States to link genomic and biomedical data, including short- and long-read whole-genome sequencing (srWGS/LRS), with rich electronic health record (EHR) information. Here, we present the first large-scale analyses of long-read sequencing (LRS) in AoU and offer a new framework for deriving genomic insights into complex structural variation (SV) of relevance to human health and disease. We performed joint analyses of 1,027 individuals self-identifying as Black or African American, sequenced to ~8x coverage with Pacific Biosciences HiFi technology and processed using cloud-native pipelines. From these LRS data we constructed a comprehensive variant callset encompassing known (FMR1 and HTT) and novel repeat expansions, clinically relevant haplotypes at loci inaccessible to srWGS, and haplotypes relevant to disease risk (HLA) and pharmacogenomics (CYP2D6), including SNVs, indels, and SVs. We developed methods for cohort-level variant calling and a scalable workflow to impute >750,000 of these SVs into existing srWGS datasets for trait association and human disease studies. Expanding to 10,000 self-identified Black or African American AoU participants with srWGS and matched EHRs, we identified 291 SV-disease associations (p < 1×10-5) spanning 226 conditions with 50.9% of associations involving SVs absent from the matched srWGS callset. Across the 226 traits, after fine-mapping using SVs and SNVs we identified 191 SV-disease pairs spanning 160 traits (70.8%) where the SV had the strongest association within the locus. Associations specific to those with computed ancestry similar to the African reference population exhibited larger effect sizes and lower allele frequencies, consistent with high-risk, ancestry-specific variants. These results demonstrate that the integration of LRS into AoU and future biobank initiatives can provide transformative new insights into genomic variation with potentially profound impact on precision medicine.