Most genetic variants associated with complex traits are hypothesized to regulate gene expression. To understand the genetics underlying gene expression variability, we characterized 14,324 RNA-sequencing samples from the Trans-Omics for Precision Medicine program and performed expression and splicing quantitative trait locus (e/sQTL) analyses in six tissues and cell types, including whole blood (n = 6454) and lung (n = 1291). We detected tens of thousands of secondary cis-e/sQTLs, showing that secondary cis-e/sQTL discovery remains unsaturated. We fine-mapped UK Biobank-derived genome-wide association study (GWAS) signals from 164 traits and identified e/sQTL colocalizations for 10,611 GWAS signals, including 7096 that colocalize with secondary e/sQTLs. Our results suggest that even larger e/sQTL analyses will uncover additional secondary e/sQTLs, further benefiting GWAS interpretation.
Large-scale multiancestry genome-wide association studies have identified hundreds of loci associated with type 2 diabetes (T2D) and glycemic traits, yet imputed genotyping arrays limit the detection of low-frequency and rare variants. Whole-genome sequencing (WGS) offers a more complete view of genetic variation, especially across diverse populations. We analyzed high-coverage (38×) WGS data from 21,913 T2D case subjects, 61,036 control subjects, and up to 50,011 individuals with no diabetes with fasting glucose, fasting insulin, and HbA1c from the National Heart, Lung, and Blood Institute Trans-Omics for Precision Medicine Program. We performed single-variant association testing, conditional analysis, fine-mapping, and Bayesian colocalization to identify genetic signals and assess regulatory relevance in diabetes-related tissues. We identified 76 distinct association signals across 34 loci, including novel variants at DUSP9 for T2D, and ROBO1, NDN, and MYT1 for HbA1c. Fine-mapping narrowed credible sets and improved causal variant resolution. Colocalization highlighted 80 expression signals in diabetes-related tissues, linking genetic associations to functional regulatory mechanisms. Our findings demonstrate the utility of WGS to uncover novel variants in diverse populations, enhance locus resolution, and link regulatory variation to disease-relevant tissues. This work refines the genetic architecture of T2D and glycemic traits and supports precision medicine efforts targeting diverse populations. ARTICLE HIGHLIGHTS:We aimed to improve understanding of the genetic architecture of type 2 diabetes and glycemic traits by leveraging whole-genome sequencing in diverse populations. Our goal was to identify novel variants, refine known loci, and link genetic signals to regulatory mechanisms through colocalization with expression quantitative trait loci. We discovered novel variants, significantly improved fine-mapping resolution, and identified 80 regulatory colocalization signals in diabetes-relevant tissues. These findings support precision medicine approaches by connecting genetic variation to functional biology in type 2 diabetes.
Summary Genotype imputation remains essential for large-scale human genetics studies, but its performance is limited by the size and ancestral diversity of available reference panels, reducing accuracy for rare variants and underrepresented populations. Here, we present a cloud-based imputation service built on a multi-ancestry reference panel derived from 515,579 jointly phased genomes from the All of Us (N=414,830) and National Human Genome Research Institute’s Analysis, Visualization, and Informatics Lab-space (AnVIL, N=100,749) datasets. The All of Us + AnVIL reference panel is highly diverse and includes 261,163 participants most genetically similar to non-European reference populations, spanning 665,398,839 high-quality autosomal sites, representing a nearly 50% increase over TOPMed, the previous largest imputation service. Across multiple ancestry groups, the panel enables accurate imputation (empirical R² > 0.8) for variants with allele frequencies as low as 0.2%, extending reliable imputation into the rare-variant frequency spectrum, including allele frequencies down to 0.002% and 0.006% for samples with European ancestry and African ancestry in the United States, respectively. Compared with TOPMed, the panel improves imputation accuracy across all ancestry groups except Africans, and recovers additional trait-associated variants not represented in existing reference panels. To facilitate broad community access while preserving participant privacy, we deploy the panel through a secure cloud-based imputation platform using privacy-preserving recombined haplotypes. This resource establishes a new foundation for genome-wide association studies (GWAS) and fine-mapping, especially in previously underrepresented populations.
Heart failure (HF) is a leading global cause of morbidity and mortality, yet the regulatory molecular mechanisms that link genetic variation to cardiac dysfunction remain elusive. To bridge this gap, we created the Trans-Omics for Precision Medicine in Congestive Heart Failure (TOPCHeF) resource, a multi-omics dataset comprising >700 human left-ventricular tissue samples, including dilated cardiomyopathy (DCM), ischemic cardiomyopathy (ICM), and non-failing controls, with paired whole-genome and RNA sequencing. By mapping expression- (eQTL) and splicing- (sQTL) quantitative trait loci directly in diseased human hearts, we identified over 10,000 transcripts with significant eQTL and 8,600 isoforms with significant sQTL, across both coding and non-coding genes, many of which overlap loci previously associated with HF and emerging novel gene associations. Single-locus colocalization with a largescale DCM genome-wide association study revealed 21 expression and 17 splicing-QTL that share causal variants with disease risk. These include known Mendelian cardiomyopathy risk genes such as FLNC and ACTN2, and novel regulatory candidates like CAMK2D, LMF1, MYOZ1, SKI, SYNPO2L, and TKT. Several loci also showed coordinated effects on both gene expression and RNA splicing, implicating calcium signaling, cytoskeletal organization, and metabolic pathways in HF pathogenesis. Together, these results help define the regulatory landscape of the failing human heart and establish TOPCHeF as a foundational resource for connecting genetic variation to transcriptional and splicing molecular mechanisms in HF research.
Chronic obstructive pulmonary disease (COPD) exhibits marked heterogeneity in lung function decline, mortality, exacerbations, and other disease-related outcomes. Omic risk scores (ORS) estimate the cumulative contribution of omics, such as the transcriptome, proteome, and metabolome, to a particular trait. This study evaluated associations between blood-based ORS and COPD-related traits in both smoking-enriched and general population cohorts. ORS were developed and tested in 3,339 participants of Genetic Epidemiology of COPD (COPDGene) with blood RNA-sequencing, proteomic, and metabolomic data. Single- and multi-omic risk scores were trained on 24 cross-sectional and five longitudinal traits using 80
Whole genome sequence (WGS) data in multi-ancestry samples supports discovery of low-frequency or population-specific genetic variants associated with chronic obstructive pulmonary disease (COPD) and lung function. We performed single variant, structural variant, and gene-based analysis of pulmonary function (FEV1, FVC and FEV1/FVC) and COPD case–control status in 44,287 multi-ancestry participants from the NHLBI Trans-Omics for Precision Medicine (TOPMed) Program. We validated findings using the UK Biobank and assessed implicated genes using lung single-cell RNA-seq (scRNA-seq) data sets. Applying a genome-wide significance threshold (P < 5 × 10–9), we replicated known loci and identified novel associations near LY86, MAGI1, GRK7, and LINC02668. Colocalization with gene expression quantitative trait loci (eQTL) from the Lung Tissue Research Consortium highlighted known candidate genes including ADAM19, THSD4, C4B, and PSMA4, which were not identified through other eQTL sources. Multi-ancestry analysis improved fine-mapping resolution (e.g., HTR4 and RIN3). Gene-based analysis identified and replicated HMCN1. In human lung scRNA-seq data sets, lung epithelial cells and immune cell types showed enriched expression, while fibroblasts showed higher expression for HMCN1. CRISPR targeting HMCN1 in IMR90 demonstrated reduced expression of collagen genes. Large-scale multi-ancestry WGS analysis improves variant discovery and fine-mapping resolution for lung function and COPD and highlights biologically relevant genes and pathways.
Rare coding genetic variants may exert large effects on risk of common disease, yet their contribution to disease architecture and their utility in gene prioritization remain limited by inadequate sample sizes. Here, we performed a massive-scale rare variant association study (RVAS), analyzing over 1.1 million sequenced participants among which 130,000 had atrial fibrillation (AF). Through a multi-mask burden testing approach, we identified 15 genes significantly associated with AF through rare large-effect variation. Integrative analyses revealed strong convergence between genes implicated by rare and common variation, and highlighted instances where RVAS data may aid in GWAS prioritization. Nevertheless, several RVAS genes were not among GWAS loci ( FAM189A2 , ACTC1 , FNIP1 , FBN1 ), or were not nominated through contemporary GWAS prioritization ( KDM5B , ZFP36L2 ). Finally, we observed that ultra-rare protein-disrupting variants - concentrated in a small number of large-effect size genes - explained at least 2% of AF susceptibility across European and African ancestry groups. These findings refine the genetic architecture of AF, while highlighting the value and cost of RVAS for genomic discovery in common disease.
Atrial fibrillation (AF) is a prevalent and morbid abnormality of the heart rhythm with a strong genetic component. Here, we meta-analyzed genome and exome sequencing data from 36 studies that included 52,416 AF cases and 277,762 controls. In burden tests of rare coding variation, we identified novel associations between AF and the genes MYBPC3, LMNA, PKP2, FAM189A2 and KDM5B. We further identified associations between AF and rare structural variants owing to deletions in CTNNA3 and duplications of GATA4. We broadly replicated our findings in independent samples from MyCode, deCODE and UK Biobank. Finally, we found that CRISPR knockout of KDM5B in stem-cell-derived atrial cardiomyocytes led to a shortening of the action potential duration and widespread transcriptomic dysregulation of genes relevant to atrial homeostasis and conduction. Our results highlight the contribution of rare coding and structural variants to AF, including genetic links between AF and cardiomyopathies, and expand our understanding of the rare variant architecture for this common arrhythmia.
The All of Us Research Program (AoU) is a national biobank seeking to enroll one million individuals in the United States to link genomic and biomedical data, including short- and long-read whole-genome sequencing (srWGS/LRS), with rich electronic health record (EHR) information. Here, we present the first large-scale analyses of long-read sequencing (LRS) in AoU and offer a new framework for deriving genomic insights into complex structural variation (SV) of relevance to human health and disease. We performed joint analyses of 1,027 individuals self-identifying as Black or African American, sequenced to ~8x coverage with Pacific Biosciences HiFi technology and processed using cloud-native pipelines. From these LRS data we constructed a comprehensive variant callset encompassing known (FMR1 and HTT) and novel repeat expansions, clinically relevant haplotypes at loci inaccessible to srWGS, and haplotypes relevant to disease risk (HLA) and pharmacogenomics (CYP2D6), including SNVs, indels, and SVs. We developed methods for cohort-level variant calling and a scalable workflow to impute >750,000 of these SVs into existing srWGS datasets for trait association and human disease studies. Expanding to 10,000 self-identified Black or African American AoU participants with srWGS and matched EHRs, we identified 291 SV-disease associations (p < 1×10-5) spanning 226 conditions with 50.9% of associations involving SVs absent from the matched srWGS callset. Across the 226 traits, after fine-mapping using SVs and SNVs we identified 191 SV-disease pairs spanning 160 traits (70.8%) where the SV had the strongest association within the locus. Associations specific to those with computed ancestry similar to the African reference population exhibited larger effect sizes and lower allele frequencies, consistent with high-risk, ancestry-specific variants. These results demonstrate that the integration of LRS into AoU and future biobank initiatives can provide transformative new insights into genomic variation with potentially profound impact on precision medicine.
and Purpose: Osteosarcoma (OS) and Leiomyosarcoma (LMS) are sarcomas with complex genomes for which there has been limited progress in identifying new treatments and improving outcomes. While slow progress is partially due to insufficient genomic characterization, generating large genomic datasets has been challenging because these are rare cancers. The OS and LMS Projects use various approaches to directly engage pediatric and adult participants with OS and LMS in genomics research. Working with patients and advocates at the design stage, we created websites (OSProject.org and LMSProject.org) where patients register and consent to participation. Any patient with OS or LMS living in the United States or Canada is eligible. Participant outreach approaches include partnership with advocacy organizations, webinars, social media posts, meeting presentations, stakeholder and physician engagement committees, and direct mailings. Blood and saliva are collected directly from consented participants by mail, archival FFPE tumor samples are obtained from pathology departments and medical records are requested from treating institutions. WES, WGS, DNA panel sequencing and RNASeq of tumor and germline (T/N) is performed. Results are shared with patient, advocacy, physician, and research communities in several ways. Individual participants receive a shared learning report describing the somatic variants identified in their tumor from T/N clinical WES and are offered clinical germline genetic testing and genetic counseling. The study outreach team has participated in or led a total of 33 online and in person events not including social media posts or stakeholder meetings. So far, in 26 months 515 LMS patients (ages 16-83y; median 55) and 145 OS patients (ages 7-74y; median 20) have consented. Thus far, 274 and 76 tumor samples and 555 and 127 germline samples have been obtained from LMS and OS consented participants, respectively. Analysis of the first 20 LMS participants with T/N WGS showed widespread alteration of TP53 (11 pts), RB1 (8 pts), and PTEN (16 pts), concordant with findings in past studies. Likewise, the first 7 OS pts with T/N WGS demonstrated inactivation of TP53 (4 pts) and RB1 (2 pts), as expected from past OS cohorts. Using community partnerships and direct outreach to connect and engage with participants and a virtual consenting process for genomics research in rare cancers is feasible. It is possible to obtain germline samples directly from about half of participants and archival tumor samples from treating pathology departments for about one third of participants consented in direct-to-patient online genomics studies. We have been able to utilize archival tumor samples to identify, with T/N WES/WGS, expected genomic events in complex genome cancers. Recruitment and sequencing are ongoing. Julia M. Wong, David Merrell, Eirian Siegal-Botti, Noorshifa Arssath, Carrie Cibulskis, Evelina Ceca, Alanna Church, Alex Wilson, Lorena Lazo De La Vega, Jill Stopfer, Ellen Sukharevsky, Nia Daley, Anusha Sharma, Sidney Benich, Zachary Kahn, Lauren Fisher, Parker Chastain, Brendan Reardon, Taisha Hendrickson, Colleen Nguyen, Melissa Chiumiento, Melissa Mirick, Noriela Elia, Priscilla Merriam, Eliezer Van Allen, Judy Garber, Riaz Gillani, Chandrajit Raut, Stacey Gabriel, Timothy Rebbeck, Jason L. Hornick, Jennifer Mack, Suzanne George, Diane Diehl, Gad Getz, Katherine A. Janeway. Directly engaging participants in rare cancer research is feasible: the osteosarcoma and leiomyosarcoma projects [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 630.
Most genetic variants associated with complex traits and diseases occur in non-coding genomic regions and are hypothesized to regulate gene expression. To understand the genetics underlying gene expression variability, we characterize 14,324 ancestrally diverse RNA-sequencing samples from the NHLBI Trans-Omics for Precision Medicine (TOPMed) program and integrate whole genome sequencing data to perform cis and trans expression and splicing quantitative trait locus (cis-/trans-e/sQTL) analyses in six tissues and cell types, most notably whole blood (N=6,454) and lung (N=1,291). We show this dataset enables greater detection of secondary cis-e/sQTL signals than was achieved in previous studies, and that secondary cis-eQTL and primary trans-eQTL signal discovery is not saturated even though eGene discovery is. Most TOPMed trans-eQTL signals colocalize with cis-e/sQTL signals, suggesting many trans signals are mediated by cis signals. We fine-map European UK BioBank GWAS signals from 164 traits and colocalize the resulting 34,107 fine-mapped GWAS signals with TOPMed e/sQTL signals, finding that of 10,611 GWAS signals with a colocalization, 7,096 GWAS signals colocalize with at least one secondary e/sQTL signal. These results demonstrate that larger e/sQTL analyses will continue to uncover secondary e/sQTL signals, and that these new signals will benefit GWAS interpretation.
In studies of individuals of primarily European genetic ancestry, common and low-frequency variants and rare coding variants have been found to be associated with the risk of bipolar disorder (BD) and schizophrenia (SZ). However, less is known for individuals of other genetic ancestries or the role of rare non-coding variants in BD and SZ risk. We performed whole-genome sequencing (∼27X) of African American individuals: 1,598 with BD, 3,295 with SZ, and 2,651 unaffected controls (InPSYght study). We increased power by incorporating 14,812 jointly called psychiatrically unscreened ancestry-matched controls from the Trans-Omics for Precision Medicine (TOPMed) Program for a total of 17,463 controls (∼37X). To identify variants and sets of variants associated with BD and/or SZ, we performed single-variant tests, gene-based tests for singleton protein truncating variants, and rare and low-frequency variant annotation-based tests with conservation and universal chromatin states and sliding windows. We found suggestive evidence of the association of BD with single variants on chromosome 18 and of lower BD risk associated with rare and low-frequency variants on chromosome 11 in a region with multiple BD genome-wide association study loci, using a sliding window approach. We also found that chromatin and conservation state tests can be used to detect differential calling of variants in controls sequenced at different centers and to assess the effectiveness of sequencing metric covariate adjustments. Our findings reinforce the need for continued whole-genome sequencing in additional samples of African American individuals and more comprehensive functional annotation of non-coding variants.
To better characterize the potential biological mechanisms underlying insulin resistance (IR) and dementia, we derive cross-population and population specific polygenic scores [PSs] for fasting insulin and IR-related partitioned PSs [pPSs]. We conduct a cross-sectional study of the associations of these genetic scores with neurological outcomes in >17k participants (36% men, mean age 55 yrs) from the Trans-Omics for Precision Medicine (TOPMed) program (50% Non-Hispanic White, 23% Black/African American, 21% Hispanic/Latino American, and 4% Asian American). We report significant negative associations (P < 0.002) of the cross-population (P = 1.3 × 10-5) and European (PEA = 3.0 × 10-8) fasting insulin PSs with total cranial volume, and of a metabolic syndrome European PS with general cognitive function (BEA = -0.13, PEA = 0.0002) and lateral ventricular volume (BEA = 0.09, PEA = 0.002). We identify suggestive negative associations (P < 0.007) of metabolic syndrome and obesity pPSs with general cognitive function, and of lipodystrophy pPSs with total cranial volume. A higher genetic predisposition to IR is associated with lower brain size, and a genetic predisposition to specific IR-related type 2 diabetes subtypes, such as metabolic syndrome and mechanisms of IR mediated through obesity and lipodystrophy, is potentially involved in cognitive decline.
Personalized cancer vaccines (PCVs) can generate circulating immune responses against predicted neoantigens1–6. However, whether such responses can target cancer driver mutations, lead to immune recognition of a patient’s tumour and result in clinical activity are largely unknown. These questions are of particular interest for patients who have tumours with a low mutational burden. Here we conducted a phase I trial (ClinicalTrials.gov identifier NCT02950766) to test a neoantigen-targeting PCV in patients with high-risk, fully resected clear cell renal cell carcinoma (RCC; stage III or IV) with or without ipilimumab administered adjacent to the vaccine. At a median follow-up of 40.2 months after surgery, none of the 9 participants enrolled in the study had a recurrence of RCC. No dose-limiting toxicities were observed. All patients generated T cell immune responses against the PCV antigens, including to RCC driver mutations in VHL, PBRM1, BAP1, KDM5C and PIK3CA. Following vaccination, there was a durable expansion of peripheral T cell clones. Moreover, T cell reactivity against autologous tumours was detected in seven out of nine patients. Our results demonstrate that neoantigen-targeting PCVs in high-risk RCC are highly immunogenic, capable of targeting key driver mutations and can induce antitumour immunity. These observations, in conjunction with the absence of recurrence in all nine vaccinated patients, highlights the promise of PCVs as effective adjuvant therapy in RCC. A phase I trial of a neoantigen-targeting personalized cancer vaccine led to durable and polyfunctional T cell responses and antitumour recognition, and was associated with no recurrence in patients with high-risk clear cell renal cell carcinoma.
Rare genetic variation provided by whole genome sequence datasets has been relatively less explored for its contributions to human traits. Meta-analysis of sequencing data offers advantages by integrating larger sample sizes from diverse cohorts, thereby increasing the likelihood of discovering novel insights into complex traits. Furthermore, emerging methods in genome-wide rare variant association testing further improve power and interpretability. Here, we conduct the largest meta-analysis of whole genome sequencing for low-density lipoprotein cholesterol (LDL-C), a therapeutic target for coronary artery disease, analyzing data from 246 K participants and integrating 1.23B variants from the UK Biobank and the Trans-Omics for Precision Medicine (TOPMed) program. We identify numerous rare coding and non-coding gene associations related to LDL-C, with replication across 86 K participants in All of Us. Our findings are based on single-variant analyses, rare coding and non-coding variant aggregation tests, and sliding window approaches. Through this comprehensive analysis, we identify 704 novel single-variant associations, 25 novel rare coding variant aggregates, 28 novel rare non-coding variant aggregates, and one novel sliding window aggregate. This study provides a meta-analysis framework for large-scale whole genome sequence association analyses from diverse population groups, yielding novel rare non-coding variant associations.
Blood lipid traits are treatable and heritable risk factors for heart disease, a leading cause of mortality worldwide. Although genome-wide association studies (GWAS) have discovered hundreds of variants associated with lipids in humans, most of the causal mechanisms of lipids remain unknown. To better understand the biological processes underlying lipid metabolism, we investigated the associations of plasma protein levels with total cholesterol (TC), triglycerides (TG), high-density lipoprotein cholesterol (HDL), and low-density lipoprotein cholesterol (LDL) in blood. We trained protein prediction models based on samples in the Multi-Ethnic Study of Atherosclerosis (MESA) and applied them to conduct proteome-wide association studies (PWAS) for lipids using the Global Lipids Genetics Consortium (GLGC) data. Of the 749 proteins tested, 42 were significantly associated with at least one lipid trait. Furthermore, we performed transcriptome-wide association studies (TWAS) for lipids using 9,714 gene expression prediction models trained on samples from peripheral blood mononuclear cells (PBMCs) in MESA and 49 tissues in the Genotype-Tissue Expression (GTEx) project. We found that although PWAS and TWAS can show different directions of associations in an individual gene, 40 out of 49 tissues showed a positive correlation between PWAS and TWAS signed p-values across all the genes, which suggests a high-level consistency between proteome-lipid associations and transcriptome-lipid associations.
Brain structural volumes are highly heritable and are linked to multiple neuropsychological outcomes, including Alzheimer's disease (AD). Genome-wide association studies have successfully identified genetic variants associated with intracranial volume (ICV), total brain volume (TBV), hippocampal volume (HV), and lateral ventricular volume (LVV). However, these studies mostly focused on common genetic variants with minor allele frequencies (MAF) > 1%, and individuals included in most of these studies were of predominantly European ancestry. Here, we performed whole-genome sequence (WGS) association studies of MRI brain volumes in 7,674 individuals of diverse race and ethnicity from the Trans-Omics for Precision Medicine (TOPMed) program. We identified novel genetic loci on chromosomes 13 and 16 near LINC00598 and CACNG3 associated with HV and TBV, respectively (lead variants rs115674829, P-value = 1.7×10-9 in pooled analysis and rs150440001, P-value = 6.6×10-9 in black participants). Both lead variant minor A alleles are rarer in white participants (MAF = 0.14% and 0.03%) and in Hispanic participants (MAF = 1.5% and 0.17%) but more common in black participants (MAF = 13% and 1.5%). Rare variant aggregated analyses identified RIPK1, a gene encoding a kinase involved in neuroinflammation and promising target for AD treatment, suggestively associated with LVV (P-value=5×10-6). This study provides new insights into the genetic correlates of brain structural volumes and illustrates the importance of leveraging WGS data and cohorts of diverse race and ethnicity to better characterize the genetic architecture of complex polygenic traits.
Bulk tissue molecular quantitative trait loci (QTLs) have been the starting point for interpreting disease-associated variants, while context-specific QTLs show particular relevance for disease. Here, we present the results of mapping interaction QTLs (iQTLs) for cell type, age, and other phenotypic variables in multi-omic, longitudinal data from blood of individuals of diverse ancestries. By modeling the interaction between genotype and estimated cell type proportions, we demonstrate that cell type iQTLs could be considered as proxies for cell type-specific QTL effects. The interpretation of age iQTLs, however, warrants caution as the moderation effect of age on the genotype and molecular phenotype association may be mediated by changes in cell type composition. Finally, we show that cell type iQTLs contribute to cell type-specific enrichment of diseases that, in combination with additional functional data, may guide future functional studies. Overall, this study highlights iQTLs to gain insights into the context-specificity of regulatory effects.
Delineation of structural variants (SVs) at sequence resolution in highly repetitive genomic regions has long been intractable. The sequence properties, origins, and functional effects of classes of genomic rearrangements such as ring chromosomes and Robertsonian translocations thus remain unknown. To resolve these complex structures, we leveraged several recent milestones in the field, including (1) the emergence of long-read sequencing, (2) the gapless telomere-to-telomere (T2T) assembly, and (3) a tool (BigClipper) to discover chromosomal rearrangements from long reads. We applied these technologies across 13 cases with ring chromosomes, Robertsonian translocations, and complex SVs that were unresolved by short reads, followed by validation using optical genome mapping (OGM). Our analyses resolved 10 of 13 cases, including a Robertsonian translocation and all ring chromosomes. Multiple breakpoints were localized to genomic regions previously recalcitrant to sequencing such as acrocentric p-arms, ribosomal DNA arrays, and telomeric repeats, and involved complex structures such as a deletion-inversion and interchromosomal dispersed duplications. We further performed methylation profiling from long-read data to discover phased differential methylation in a gene promoter proximal to a ring fusion, suggesting a long-range position effect (LRPE) with heterochromatin spreading. Breakpoint sequences suggested mechanisms of SV formation such as microhomology-mediated and non-homologous end-joining, as well as non-allelic homologous recombination. These methods provide some of the first glimpses into the sequence resolution of Robertsonian translocations and illuminate the structural diversity of ring chromosomes and complex chromosomal rearrangements with implications for genome biology, prediction of LRPEs from integrated multi-omics technologies, and molecular diagnostics in rare disease cases.