Abstract Background Native long-read DNA sequencing simultaneously captures genetic variants and epigenetic modifications from single molecules, but preserving molecular length and base modifications currently depends on cold-chain infrastructure that limits access to well-resourced settings. Results We demonstrate that ensilication, the encapsulation of DNA within silica matrices, preserves DNA at ambient temperature for 30 days with sequencing performance equivalent to conventional − 80 °C freezing. Across three Genome-in-a-Bottle reference genomes, ensilicated and frozen samples show no significant differences in read length (N50 ~ 8,000–11,000 bp), variant-calling accuracy, or genome-wide CpG methylation. Single-read methylation calls benchmarked against an independent bisulfite-sequencing reference confirm that ensilication introduces no detectable bias, with per-read accuracy differing by less than 0.4% between preservation conditions. Ensilicated DNA tolerates repeated handling better than frozen samples and maintains fragment integrity under accelerated weathering. In two patients with rare genetic disorders, ambient-preserved DNA resolves a de novo variant in the segmentally duplicated GTF2I locus and detects methylation patterns consistent with KDM2A-related disorder. Conclusions Ensilication enables diagnostic-quality native long-read sequencing without cold-chain infrastructure, supporting ambient storage and transport while preserving both sequence and methylation information.
Structural variants (SVs) are a major source of genetic variation yet remain underexplored in healthy aging and neurodegenerative diseases. We performed nanopore long-read genome sequencing (lrGS) on 551 deeply-phenotyped individuals from Stanford's Aging and Memory Study and Alzheimer's Disease Research Center, generating a comprehensive SV map integrated with matched methylation, transcriptomic, and proteomic data. Over 60% of SVs identified by lrGS were not detected with short-read WGS, including many poorly tagged by single-nucleotide variants (SNVs). We discovered >60,000 SV-QTLs across molecular traits and showed that SVs were more likely than SNVs to be fine-mapped as causal. Colocalization with Alzheimer's and Parkinson's disease GWAS implicated SVs at multiple loci, including TMEM106B, BIN3, and NBEAL1. Multi-omic outlier enrichment and Bayesian modeling prioritized rare functional SVs near known risk genes. Combined, these data reveal widespread regulatory SVs in healthy aging and neurodegeneration, underscoring the importance of lrGS in deciphering complex genetic architecture.
Rare structural variants (SVs)-insertions, deletions, and complex rearrangements-can cause Mendelian disease, yet they remain difficult to accurately detect and interpret. We sequenced and analyzed Oxford Nanopore Technologies long-read genomes of 68 individuals from the undiagnosed disease network (UDN) with no previously identified diagnostic mutations from short-read sequencing. Using our optimized SV detection pipelines and 571 control long-read genomes, we detected 716 long-read rare (MAF < 0.01) SV alleles per genome on average, achieving a 2.4× increase from short reads. To characterize the functional effects of rare SVs, we assessed their relationship with gene expression from blood or fibroblasts from the same individuals and found that rare SVs overlapping enhancers were enriched (LOR = 0.46) near expression outliers. We also evaluated tandem repeat expansions (TREs) and found 14 rare TREs per genome; notably, these TREs were also enriched near overexpression outliers. To prioritize candidate functional SVs, we developed Watershed-SV, a probabilistic model that integrates expression data with SV-specific genomic annotations, which significantly outperforms baseline models that do not incorporate expression data. Watershed-SV identified a median of eight high-confidence functional SVs per UDN genome. Notably, this included compound heterozygous deletions in FAM177A1 shared by two siblings, which were likely causal for a rare neurodevelopmental disorder. Our observations demonstrate the promise of integrating long-read sequencing with gene expression toward improving the prioritization of functional SVs and TREs in rare disease patients.
Alzheimer’s disease (AD) is the most common form of dementia. Neuropathologically, AD stands out as a mixed proteinopathy. Beta-amyloid and tau biomarkers can now add in-vivo support to the AD diagnosis. Rarely, a patient with AD confirmed at autopsy may have a negative amyloid PET. Here we describe a pedigree in whom the index case was amyloid PET-negative. We performed genetic sequencing of two affected siblings to identify a set of candidate single nucleotide variants (SNVs) and structural variants (SVs) associated with disease and rare in healthy older controls from several large genetic databases. We performed long-read sequencing (LRS) to comprehensively evaluate both SNVs and SVs. The index case, an APOE3/E4 male, developed memory trouble at 68 that progressed to probable AD at 72. Notably (Fig. 1) his amyloid-PET scan, which was confirmed to be of high technical quality, was negative. His tau PET scan was positive and his plasma Abeta42/40 was in the expected range for Stanford AD patients. His family history is extensive and suggestive of an autosomal dominant pattern of late-onset AD (Fig. 2). The patient’s sister was diagnosed with AD at 62. We filtered their shared heterozygous SNVs keeping those with a minor allele frequency < 0.01 in gnomAD and that were not present in any ADSP healthy controls (HC) over 70y.o. (N = 19771;62.1%females;81±6.4y.o). Shared SVs were kept as candidates if they were not seen in any Stanford LRS HC (N = 95;59%females;72±1.5y.o.) Neuropathological analysis revealed that both siblings met the A3B3C3 criteria for AD and had cerebral amyloid angiopathy. Among rare mutations on genes expressed in the brain (Table 1), a novel missense in ADNP (Activity-Dependent-Neuroprotector-Homeobox), and a 267bp deletion on TET1 (Tet-Methylcytosine-Dioxygenase) were identified. The 3426bp insertion on ATP8A2 (ATPase-Phospholipid-Transporting-8A2) was the longest, rare insertion found. Additionally, a large duplication (∼20Kbp) and inversion (∼160Kbp) were observed. This study highlights an unusual AD pedigree and emphasizes the potential utility of LRS in identifying causal SNVs and SVs. In future work we plan to perform additional amyloid stains to understand why the index case PET was amyloid-negative. We are also assessing additional family members to improve our ability to identify a causal mutation.
Long-read DNA sequencing detects genetic variants and epigenetic modifications simultaneously, but optimal preservation of molecular length and base modifications typically requires cold-chain infrastructure. This infrastructure dependence restricts genomics to well-resourced laboratories and limits field research. We show that ensilication, the encapsulation of DNA within silica matrices, preserves DNA integrity at ambient temperature equivalent to -80 °C freezing. Using Genome-in-a-Bottle (GIAB) reference samples, we show that ensilicated DNA maintains sequencing performance comparable to frozen samples, while exhibiting resistance to degradation during repeated handling and accelerated weathering conditions simulating decades of ambient storage.. In real world samples from patients, native long-read sequencing of ensilicated samples resolved a de novo variant in the segmentally duplicated GTF2I locus in one case and detected methylation episignatures diagnostic of KDM2A-related disorder in another. Ensilication removes cold-chain dependence for diagnostic-quality long-read sequencing, expanding access to molecular diagnostics in resource-limited settings.
Although the gross morphology of the heart is conserved across mammals, subtle interspecific variations exist in the cardiac phenotype, which may reflect evolutionary divergence among closely-related species. Here, we compare the left ventricle (LV) across all extant members of the Hominidae taxon, using 2D echocardiography, to gain insight into the evolution of the human heart. We present compelling evidence that the human LV has diverged away from a more trabeculated phenotype present in all other great apes, towards a ventricular wall with proportionally greater compact myocardium, which was corroborated by post-mortem chimpanzee (Pan troglodytes) hearts. Speckle-tracking echocardiographic analyses identified a negative curvilinear relationship between the degree of trabeculation and LV systolic twist, revealing lower rotational mechanics in the trabeculated non-human great ape LV. This divergent evolution of the human heart may have facilitated the augmentation of cardiac output to support the metabolic and thermoregulatory demands of the human ecological niche.
Single nucleotide variants (SNVs) near TMEM106B have been associated with risk of frontotemporal lobar dementia with TDP pathology (FTLD-TDP) but the causal variant at this locus has not yet been isolated. The initial leading FTLD-TDP genome-wide association study (GWAS) hit at this locus, rs1990622, is intergenic and is in linkage disequilibrium (LD) with a TMEM106B coding SNV, rs3173615. We developed a long-read sequencing (LRS) dataset of 407 individuals in order to identify structural variants associated with neurodegenerative disorders. We identified a prevalent 322 base pair deletion on the TMEM106B 3 'untranslated region (UTR) that was in perfect linkage with rs1990622 and near-perfect linkage with rs3173615 (genotype discordance in two of 274 individuals who had LRS and short-read next-generation sequencing). In Alzheimer's Disease Sequencing Project (ADSP) participants, this deletion was in greater LD with rs1990622 (R2=0.920916, D'=0.963472) than with rs3173615 (R2=0.883776, D'=0.963575). rs1990622 and rs3173615 are less closely linked (R2=0.7403, D'=0.9915) in African populations. Among African ancestry individuals in the ADSP, the deletion is in even greater LD with rs1990622 (R2=0.936841, D'=0.976782) than with rs3173615 (R2=0.764242, D'=0.974406). Querying publicly available genetic datasets with associated mRNA expression and protein levels, we confirmed that rs1990622 is consistently a protein quantitative trait locus but not an expression quantitative trait locus, consistent with a causal variant present on the TMEM106B 3'UTR. In summary, the TMEM106B 3' UTR deletion is a large genetic variant on the TMEM106B transcript that is in higher LD with the leading GWAS hit rs1990622 than rs3173615 and may mediate the protective effect of this locus in neurodegenerative disease.
Long read sequencing offers benefits for the detection of structural variation in Mendelian disease. Here, we applied a new technology that generates contiguous long reads via tagmentation and sequencing by synthesis to a small cohort of patients with undiagnosed disease from the Undiagnosed Diseases Network. We first compare sequencing from the HG002 benchmark sample from Genome In A Bottle using nanopore sequencing (R10.4.1, duplex reads, Oxford Nanopore), single molecule real time sequencing (Revio SMRT cell, Pacific Biosciences) and complete long read sequencing (S4 flowcell, Novaseq, Illumina). Coverage was 33-35x across platforms. Read length N50 was 6.5kb (ICLR), 16.9kb (SMRT), and 33.8kb (ONT). We noted small differences in single nucleotide variant F1 scores across long read technologies with single nucleotide variant F1 scores (0.985-0.999) exceeding indel scores (0.78-0.99) and structural variant scores (0.74-0.96). We applied CLR sequencing to seven undiagnosed patients. In one patient, we detected and prioritized a novel 16kb intragenic duplication encompassing exons 5 and 6 in EHMT1. Resolution of the breakpoints and examination of flanking sequences revealed that the duplication was present in tandem and was predicted to result in a frameshift of the amino acid sequence and an early termination codon. It resulted in a diagnosis of Kleefstra syndrome. The variant was confirmed with targeted EHMT1 clinical testing and detected via nanopore and SMRT sequencing. In summary, we report the early clinical application of complete long read sequencing to a small cohort of undiagnosed patients. ### Competing Interest Statement Euan A. Ashley has received support in kind from Illumina, PacBio, and Nanopore. ### Funding Statement This study was funded by the Undiagnosed Diseases Network. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Participants sequenced in this study were enrolled in the Undiagnosed Diseases Network (UDN) at Stanford Medicine and provided informed consent. The study was granted ethical approved by the central IRB at the National Institutes of Health. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes The data on HG002 Genome in a Bottle sample is available publicly and referenced in the Methods section. The Ilumina complete long read sequencing data will be available in dbGaP in accordance with Undiagnosed Diseases Network data sharing policies. As data is deposited from the Undiagnosed Disease Network Data Management Coordinating Center Gateway Database to dbGaP, the complete long read sequencing data is not immediately available. Aggregate short read genome sequencing data on all UDN participants can be accessed via dbGaP Study Accession: phs001232.v5.p2.
The ε4 allele of apolipoprotein E (APOE) is the strongest genetic risk factor for sporadic Alzheimer's Disease (AD). Knockdown of this allele may provide a therapeutic strategy for AD, but the effect of APOE loss-of-function (LoF) on AD pathogenesis is unknown. We searched for APOE LoF variants in a large cohort of older controls and patients with AD and identified six heterozygote carriers of APOE LoF variants. Five carriers were controls (ages 71-90) and one was an AD case with an unremarkable age-at-onset between 75-79. Two APOE ε3/ε4 controls (Subjects 1 and 2) carried a stop-gain affecting the ε4 allele. Subject 1 was cognitively normal at 90+ and had no neuritic plaques at autopsy. Subject 2 was cognitively healthy within the age range 75-79 and underwent lumbar puncture at between ages 75-79 with normal levels of amyloid. The results provide the strongest human genetics evidence yet available suggesting that ε4 drives AD risk through a gain of abnormal function and support knockdown of APOE ε4 or its protein product as a viable therapeutic option.
Long-read sequencing technology has enabled variant detection in difficult-to-map regions of the genome and enabled rapid genetic diagnosis in clinical settings. Rapidly evolving third-generation sequencing platforms like Pacific Biosciences (PacBio) and Oxford nanopore technologies (ONT) are introducing newer platforms and data types. It has been demonstrated that variant calling methods based on deep neural networks can use local haplotyping information with long-reads to improve the genotyping accuracy. However, using local haplotype information creates an overhead as variant calling needs to be performed multiple times which ultimately makes it difficult to extend to new data types and platforms as they get introduced. In this work, we have developed a local haplotype approximate method that enables state-of-the-art variant calling performance with multiple sequencing platforms including PacBio Revio system, ONT R10.4 simplex and duplex data. This addition of local haplotype approximation makes DeepVariant a universal variant calling solution for long-read sequencing platforms.
The sustainability of zoo populations is dependent on maintaining genetic diversity and controlling heritable disease. Here, we explore the integration of whole genome sequencing data in the management of the international zoological population of western lowland gorillas, focusing on genetic diversity and heritable diseases. By comparing kinship values derived from classical pedigree mapping and whole genome sequencing, we demonstrate that genomic data provides a more sensitive measure of relatedness. Our analysis reveals a decrease in genetic diversity due to closed breeding, emphasizing the potential for genetic intervention to mitigate negative impact on population fitness. We identify contributing factors to the decreasing genetic diversity including breeding within a closed population, unknown kinship among potential mates, and disproportionate genetic contributions from individual founders. Additionally, we highlight idiopathic myocardial fibrosis (IMF), a common cardiovascular pathology observed in zoologically housed gorillas, and identify a novel genetic variant in the TNNI3K gene that appears to be associated with this condition. These findings underscore the importance of incorporating molecular data into ex-situ population management strategies, and advocate for the adoption of advanced genomic techniques to optimize the genetic health and diversity of zoologically housed western lowland gorillas.
Alzheimer’s disease (AD) is up to 60-80% heritable, but less than ∼20% is explained by studies analyzing single nucleotide variants (SNVs). One limitation of short-read whole-genome sequencing (SRS) is the standard read length of 150 base pairs, which does not enable the detection of longer structural variants (SVs). Here we sequenced individuals using long-read sequencing (LRS), which can sequence reads with an average length of ∼20 kilobases, allowing us to identify SVs previously uncaptured. All participants underwent whole-genome LRS (∼15x coverage) and a subgroup (84%) underwent SRS. Out of 576 participants (47% males, age = 70.6±7.9 y.o.), 115 were diagnosed with AD or mild cognitive impairment, 365 were healthy controls, and 96 were diagnosed with a synucleinopathy (either Parkinson or Lewy Body disease). Eighty-three index SNVs from loci associated with AD risk through GWAS (Bellenguez et al., Nature Genetics 2022) were genotyped with 30x coverage SRS, in addition to APOE2-4. A 1Mbp window was defined around these SNVs to construct the discovery range for SVs (Sniffles2 population mode). Linkage disequilibrium (LD) was assessed between SNVs and SVs (CubeX). A total of 14854 SVs were found across the AD risk loci (Figure1). After LD calculation, N = 197 SVs had a R 2 >0.1 (Figure2). The SVs with the highest LD (R 2 >0.7) are reported in Table1, among which there is a 322 bp deletion in the 3’ UTR region of the TMEM106B in high LD (R 2 = 0.918) with the intronic SNV rs13237518. This TMEM106B locus has been previously associated with the risk of frontotemporal lobar dementia with TDP-43 pathology in addition to AD, but the causative SNV has not been identified yet. This large deletion may mediate TMEM106B ’s risk-modulating role in AD and FTLD-TDP (Chemparathy et al. medRxiv 2023). At the complex MAPT locus, three SVs showed a high LD with rs199515. Insights into the role of SVs in neurodegenerative disorders have been hampered due to limitations with SRS. Using LRS in a large AD-related sample, we characterized for the first time the genetic variation of SVs in known AD risk loci and provide a roadmap to identify potential causal SVs driving the AD association signal.
Incorporating comprehensive genetic testing such as whole genome sequencing (WGS) into the traditional diagnostic process allows clinical teams to understand a patient’s disease at its most underlying molecular basis. With rapid-WGS becoming more accessible and affordable, a typical clinic/hospital will be able to perform it efficiently and at scale, on demand for a patient. Along with improvements in sequencing platforms, innovations in computer architecture and custom hardware are the key in realizing this vision. Their combined prowess is, and will continue to be, critical in our ability to generate, transfer, analyze, curate, and infer genomic data in real time.
To the Editor: Rapid genetic diagnosis can guide clinical management, improve prognosis, and reduce costs in critically ill patients.1,2 Although most critical care decisions must be made in hours, traditional testing requires weeks and rapid testing requires days. We have found that nanopore genome sequencing can accurately and rapidly provide genetic diagnoses. Our workflow combines streamlined preparation of commercial nanopore sequencing, distributed Cloudbased bioinformatics, and a custom variant-prioritization approach (Fig. 1).3 Between December 2020 and May 2021, at two hospitals in Stanford, California, we enrolled 12 patients who were generally representative of persons living in the United States with respect to race, ethnic group, and sex (Tables S1 and S2 in the Supplementary Appendix, available with the full text of this letter at NEJM.org). We obtained an initial genetic diagnosis in 5 of the patients (Table S3). The shortest time from arrival of the blood sample in the laboratory to the initial diagnosis was 7 hours 18 minutes. After establishing a diagnosis in Patient 1, we updated our bioinformatics framework to permit the transfer of terabytes of raw signal data to Cloud storage in real time and distributed the data across multiple Cloud computing machines to achieve near real-time base calling and alignment, a step that reduced the postsequencing run time (base calling through alignment) by 93%, from 7 hours 21 minutes to 34 minutes (the average of postsequencing run times for Patients 2 to 12) (Table S5). Flow cells were washed and reused until exhaustion to reduce the sequencing cost per sample. Libraries were bar-coded in Patients 1 through 7 to prevent carryover from one sample to the next. After processing the sample obtained from Patient 7, we benchmarked and adopted a bar-code–free method to rapidly generate genome sequences.3 Removing the bar-coding process accelerated sample preparation by 37 minutes, to an average of 2.5 hours, and enabled us to load a greater amount of patients’ DNA into each flow cell (333 ng vs. 155 ng) and increase pore occupancy (to 82% from 64%) (Figs. S1 and S2 and Table S4). Our sequencing workflow generated 173 to 236 Gb of data per genome using 48 flow cells, with an alignment identity of 94% (Fig. S3) and 46 to 64× autosomal coverage (i.e., each base of each autosome was represented in 46 to 64 sequence reads) (Fig. S4). Half the sequencing throughput was in reads that were 25 kb or longer (Table S6). Small variants and structural variants were called after the reads were aligned to the GRCh37 human reference genome, which generated a median of 4,490,490 single-nucleotide variants and small insertions and deletions (indels).4,5 Custom filtration and prioritization of variants with an ultrarapid scoring system (Fig. S5) substantially decreased the number of candidate variants for manual review to a median of 29 (range, 16 to 53) for small variants and 22 (range, 11 to 37) for structural variants (Table S2). Each initial diagnosis was immediately reviewed by study and bedside physicians, and a consensus was reached as to whether the proposed variant represented the primary cause of the patient’s presentation. Diagnostic variants were identified in 5 of the 12 patients, who ranged in age from 3 months to 57 years. The findings were immediately confirmed by a laboratory certified by the Clinical Laboratory Improvement Amendments (CLIA) process and informed clinical management (including sympathectomy, heart transplantation, screening, and changes in medication) for each of the 5 patients or their family members. In one patient, a 3-month-old full-term infant who presented in status epilepticus, seizure semiology included right gaze deviation with bilateral upper-extremity clonic jerking and perioral myoclonic twitching. Interictal electroencephalography revealed abundant predominantly posterior
HomeCirculation: Genomic and Precision MedicineVol. 15, No. 2Ultra-Rapid Nanopore Whole Genome Genetic Diagnosis of Dilated Cardiomyopathy in an Adolescent With Cardiogenic Shock Free AccessLetterPDF/EPUBAboutView PDFView EPUBSections ToolsAdd to favoritesDownload citationsTrack citationsPermissions ShareShare onFacebookTwitterLinked InMendeleyReddit Jump toFree AccessLetterPDF/EPUBUltra-Rapid Nanopore Whole Genome Genetic Diagnosis of Dilated Cardiomyopathy in an Adolescent With Cardiogenic Shock John E. Gorzynski, DVM, PhD, Sneha D. Goenka, MTech, Kishwar Shafin, BS, Tanner D. Jensen, BS, Dianna G. Fisk, PhD, Megan E. Grove, MS, Elizabeth Spiteri, PhD, Trevor Pesout, BS, Jean Monlong, PhD, Jonathan A. Bernstein, MD, PhD, Scott Ceresnak, MD, Pi-Chuan Chang, PhD, Jeffrey W. Christle, PhD, Henry Chubb, MBBS, PhD, Kyla Dunn, MS, Daniel R. Garalde, PhD, Joseph Guillory, MS, Maura R.Z. Ruzhnikov, MD, Chris Wright, DPhil, Courtney J. Wusthoff, MD, Katherine Xiong, MD, Seth A. Hollander, MD, Gerald J. Berry, MD, Miten Jain, PhD, Fritz J. Sedlazeck, PhD, Andrew Carroll, PhD, Benedict Paten, PhD and Euan A. Ashley, MB, ChB, DPhil John E. GorzynskiJohn E. Gorzynski https://orcid.org/0000-0002-9034-9016 Stanford University, CA (J.E.G., S.D.G., T.D.J., E.S., J.A.B., S.C., J.W.C., H.C., M.R.Z.R., C.J.W., K.X., S.A.H., G.J.B., E.A.A.). , Sneha D. GoenkaSneha D. Goenka https://orcid.org/0000-0002-1716-7769 Stanford University, CA (J.E.G., S.D.G., T.D.J., E.S., J.A.B., S.C., J.W.C., H.C., M.R.Z.R., C.J.W., K.X., S.A.H., G.J.B., E.A.A.). , Kishwar ShafinKishwar Shafin https://orcid.org/0000-0001-5252-3434 University of California at Santa Cruz Genomics Institute, Santa Cruz, CA (K.S., T.P., J.M., M.J., B.P.). , Tanner D. JensenTanner D. Jensen Stanford University, CA (J.E.G., S.D.G., T.D.J., E.S., J.A.B., S.C., J.W.C., H.C., M.R.Z.R., C.J.W., K.X., S.A.H., G.J.B., E.A.A.). , Dianna G. FiskDianna G. Fisk Stanford Health Care, Palo Alto, CA (D.G.F., M.E.G.). , Megan E. GroveMegan E. Grove https://orcid.org/0000-0002-7972-0005 Stanford Health Care, Palo Alto, CA (D.G.F., M.E.G.). , Elizabeth SpiteriElizabeth Spiteri Stanford University, CA (J.E.G., S.D.G., T.D.J., E.S., J.A.B., S.C., J.W.C., H.C., M.R.Z.R., C.J.W., K.X., S.A.H., G.J.B., E.A.A.). , Trevor PesoutTrevor Pesout University of California at Santa Cruz Genomics Institute, Santa Cruz, CA (K.S., T.P., J.M., M.J., B.P.). , Jean MonlongJean Monlong https://orcid.org/0000-0002-9737-5516 University of California at Santa Cruz Genomics Institute, Santa Cruz, CA (K.S., T.P., J.M., M.J., B.P.). , Jonathan A. BernsteinJonathan A. Bernstein https://orcid.org/0000-0001-5369-346X Stanford University, CA (J.E.G., S.D.G., T.D.J., E.S., J.A.B., S.C., J.W.C., H.C., M.R.Z.R., C.J.W., K.X., S.A.H., G.J.B., E.A.A.). , Scott CeresnakScott Ceresnak https://orcid.org/0000-0002-9473-0105 Stanford University, CA (J.E.G., S.D.G., T.D.J., E.S., J.A.B., S.C., J.W.C., H.C., M.R.Z.R., C.J.W., K.X., S.A.H., G.J.B., E.A.A.). , Pi-Chuan ChangPi-Chuan Chang https://orcid.org/0000-0003-3021-6446 Google Inc, Mountain View, CA (P.-C.C., A.C.). , Jeffrey W. ChristleJeffrey W. Christle Stanford University, CA (J.E.G., S.D.G., T.D.J., E.S., J.A.B., S.C., J.W.C., H.C., M.R.Z.R., C.J.W., K.X., S.A.H., G.J.B., E.A.A.). , Henry ChubbHenry Chubb https://orcid.org/0000-0002-2859-3536 Stanford University, CA (J.E.G., S.D.G., T.D.J., E.S., J.A.B., S.C., J.W.C., H.C., M.R.Z.R., C.J.W., K.X., S.A.H., G.J.B., E.A.A.). , Kyla DunnKyla Dunn Stanford Children’s Health, Palo Alto, CA (K.D.). , Daniel R. GaraldeDaniel R. Garalde Oxford Nanopore Technologies, United Kingdom (D.R.G., J.G., C.W.). , Joseph GuilloryJoseph Guillory Oxford Nanopore Technologies, United Kingdom (D.R.G., J.G., C.W.). , Maura R.Z. RuzhnikovMaura R.Z. Ruzhnikov https://orcid.org/0000-0001-9610-1612 Stanford University, CA (J.E.G., S.D.G., T.D.J., E.S., J.A.B., S.C., J.W.C., H.C., M.R.Z.R., C.J.W., K.X., S.A.H., G.J.B., E.A.A.). , Chris WrightChris Wright Oxford Nanopore Technologies, United Kingdom (D.R.G., J.G., C.W.). , Courtney J. WusthoffCourtney J. Wusthoff https://orcid.org/0000-0002-1882-5567 Stanford University, CA (J.E.G., S.D.G., T.D.J., E.S., J.A.B., S.C., J.W.C., H.C., M.R.Z.R., C.J.W., K.X., S.A.H., G.J.B., E.A.A.). , Katherine XiongKatherine Xiong Stanford University, CA (J.E.G., S.D.G., T.D.J., E.S., J.A.B., S.C., J.W.C., H.C., M.R.Z.R., C.J.W., K.X., S.A.H., G.J.B., E.A.A.). , Seth A. HollanderSeth A. Hollander Stanford University, CA (J.E.G., S.D.G., T.D.J., E.S., J.A.B., S.C., J.W.C., H.C., M.R.Z.R., C.J.W., K.X., S.A.H., G.J.B., E.A.A.). , Gerald J. BerryGerald J. Berry https://orcid.org/0000-0002-6176-2629 Stanford University, CA (J.E.G., S.D.G., T.D.J., E.S., J.A.B., S.C., J.W.C., H.C., M.R.Z.R., C.J.W., K.X., S.A.H., G.J.B., E.A.A.). , Miten JainMiten Jain University of California at Santa Cruz Genomics Institute, Santa Cruz, CA (K.S., T.P., J.M., M.J., B.P.). , Fritz J. SedlazeckFritz J. Sedlazeck https://orcid.org/0000-0001-6040-2691 Baylor College of Medicine, Houston, TX (F.J.S.). , Andrew CarrollAndrew Carroll https://orcid.org/0000-0002-4824-6689 Google Inc, Mountain View, CA (P.-C.C., A.C.). , Benedict PatenBenedict Paten University of California at Santa Cruz Genomics Institute, Santa Cruz, CA (K.S., T.P., J.M., M.J., B.P.). and Euan A. AshleyEuan A. Ashley Correspondence to: Euan A. Ashley, MB, ChB, DPhil, Stanford University, 300 Pasteur Dr, Falk CVRC, CV-267, Stanford, CA 94305. Email E-mail Address: [email protected] https://orcid.org/0000-0001-9418-9577 Stanford University, CA (J.E.G., S.D.G., T.D.J., E.S., J.A.B., S.C., J.W.C., H.C., M.R.Z.R., C.J.W., K.X., S.A.H., G.J.B., E.A.A.). Originally published8 Feb 2022https://doi.org/10.1161/CIRCGEN.121.003591Circulation: Genomic and Precision Medicine. 2022;15Other version(s) of this articleYou are viewing the most recent version of this article. Previous versions: February 8, 2022: Ahead of Print Rapid genetic diagnosis has the potential to guide clinical treatment in critically ill patients leading to improved prognosis and decreased health care costs.1 Until recently, the turnaround time for whole genome diagnostic testing precluded its integration into critical care decision making (typical rapid whole genome sequencing clinical testing returns results in 5–7 days). Here, we describe a case of a teenager presenting with cardiogenic shock in whom a genetic diagnosis was made in under 12 hours using a new ultra-rapid long read whole genome sequencing assay and workflow.2,3A 13-year-old male previously in good health presented to his primary care provider with a nocturnal dry cough, decreased appetite, intermittent chest pain, and fatigue. Thoracic radiographs showed cardiomegaly leading to echocardiography, which revealed a dilated left ventricle with an ejection fraction of 29%. The patient was then transferred to Stanford Children’s Health, Lucile Packard Children’s Hospital and subsequent echocardiography revealed an ejection fraction of 20.9% and a left ventricular end-diastolic diameter of 6.6 cm (Figure [A]). Worsening end organ perfusion was observed and shortly thereafter the patient was cannulated for veno-arterial extracorporeal membrane oxygenation.Download figureDownload PowerPointFigure. Ultra-rapid genome sequencing of a patient with dilated cardiomyopathy identifies a variant in cardiac troponin. A, Baseline echocardiography (previous to extracorporeal membrane oxygenation [ECMO] cannulation). A short axis and 4 chamber view demonstrate dilated cardiomyopathy including biventricular remodeling with dilation and wall thinning, as well as mitral regurgitation. B, Right ventricular (RV) septal biopsy obtained via catheterization via the right femoral vein. Hematoxylin and eosin and trichrome stains show enlarged, irregularly shaped myocyte nuclei indicative of hypertrophy (H&E ×200; trichrome ×200). The trichrome stain highlights interstitial fibrosis indicated by the asterisk (*). C, The variant filtering and prioritization scheme allowed rapid review and interpretation of likely actionable variants by setting a threshold of 4 or higher for review and elevating the highest scoring variants for focused interpretation. This method identified a 3 base pair insertion in a short tandem repeat element of TNNT2 (ClinVar Accession: SCV002037155). D, Sanger Sequencing confirms presence of a heterozygous 3 base pair insertion in a short tandem repeat element of TNNT2. gnomAD indicates genome aggregation database; HGMD, human gene mutation database; LV, left ventricle; MAF, minor allele frequency; and RVIS, residual variation intolerance score.The differential diagnosis included lymphocytic or giant cell myocarditis, toxic cardiomyopathy, and genetic cardiomyopathy. While acute myocarditis patients often recover without the need for advanced therapies, genetic causes are typically associated with progressive disease requiring surgical mechanical support and transplantation.4,5 While toxic cardiomyopathy is a diagnosis of exclusion and myocarditis is either a diagnosis of exclusion or one secured by invasive cardiac biopsy, genetic cardiomyopathy can be conclusively demonstrated if a pathogenic variant is found in a known disease causing gene. Due to the acute onset and rapid progression of cardiogenic shock, there was a pressing need to differentiate the cause of disease. Advanced imaging (magnetic resonance imaging or F-fluorodeoxyglucose Computed Tomography-Positron Emission Tomography) can provide evidence for inflammatory cardiomyopathy, however, the patient’s condition limited the possibility of this testing. As such, the patient underwent right ventricular cardiac biopsy via catheterization of the right femoral vein (Figure [B]). In addition, the patient was enrolled in the ultra-rapid whole genome sequencing research program.Two milliliters of whole blood in ethylenediaminetetraacetic (EDTA) acid was collected and genomic DNA extracted using a modified Puregene (Qiagen) method, followed by sequencing library preparation using LSK-109 and native barcodes from Oxford Nanopore Technologies. Sample processing took a total of 4 hours and 12 minutes. The sequencing library was distributed over 48 PromethION flow cells which sequenced simultaneously for a total of 2 hours and 42 minutes resulting in 204 gigabases of sequencing reads with an N50 of 22 kilobases. In real time, the raw sequencing data was uploaded to a cloud server, where base calling and alignment occurred in parallel to sequencing. Small variants (single-nucleotide variants and small insertions/deletions) were called using PEPPER-Margin-DeepVariant resulting in identification of 4 371 501 genomic variants.6 Structural variants were called using Sniffles resulting in 20 prioritized structural variants.7 Sequencing data analysis and preparation took a total of 3 hours 36 minutes postsequencing. Variant call files were then transferred to the curation team where filtering, prioritization and manual curation identified a heterozygous duplication in a short tandem repeat element in TNNT2 (487_489dup GAG), an integral component of the cardiac sarcomere and a gene known to be associated with dilated cardiomyopathy (Figure [C]). Curation time took a total of 46 minutes, identifying a candidate variant in 11 hours and 16 minutes after the sample preparation began. Sanger sequencing confirmed the presence of this variant (Figure [D]) and parental testing later revealed it to be de novo, confirming this variant as likely pathogenic.The right ventricular biopsy produced 3 endomyocardial samples showing myocyte hypertrophy and patchy interstitial fibrosis compatible with dilated cardiomyopathy. No myocarditis, infiltrative disorders or metabolic alterations were present.Ultra-rapid whole genome sequencing identified a likely pathogenic variant in TNNT2 while pathology findings provided no evidence for an inflammatory cause, supporting the diagnosis of dilated cardiomyopathy with a genetic cause. These findings were available before the discussion of transplant listing. In contrast, a clinical panel sent to a commercial laboratory did not return results with the TNNT2 variant until the transplant listing decision was made, emphasizing the impact of rapid turnaround testing. The patient received a heart transplant 21 days after listing.Article InformationSources of FundingThis work was supported by in-kind contributions from Oxford Nanopore, Google, and Nvidia. University California Santa Cruz Genomics Institute and Stanford University unrestricted funds financially supported this study.Disclosures K. Shafin has performed paid internships at NVIDIA Corp and Google LLC, and presented a talk at an Oxford Nanopore Technologies (ONT) sponsored event. Drs Chang and Carroll are employees of Google LLC and own Alphabet stock as part of the standard compensation package. Dr Garalde, J. Guillory, and Dr Wusthoff are employees of ONT and share/share option holders. Dr Jain has received reimbursement for travel, accommodation and conference fees to speak at events organized by ONT. Dr Sedlazeck received travel compensation from Pacific Biotechnology and ONT. Dr Ashley is cofounder of Personalis, Deepcell, and Svexa, Advisor to Apple, and a Non-Executive Director of AstraZeneca. The other authors report no conflicts. Google employees did not have access to patient data.FootnotesFor Sources of Funding and Disclosures, see page 169.Correspondence to: Euan A. Ashley, MB, ChB, DPhil, Stanford University, 300 Pasteur Dr, Falk CVRC, CV-267, Stanford, CA 94305. Email [email protected]eduReferences1. Buchan JG, White S, Joshi R, Ashley EA. Rapid genome sequencing in the critically ill.Clin Chem. 2019; 65:723–726. doi: 10.1373/clinchem.2018.293506CrossrefMedlineGoogle Scholar2. Gorzynski JE, Goenka SD, Shafin K, Jensen TD, Fisk DG, Grove ME, Spiteri E, Pesout T, Monlong J, Baid G, et al.. Ultrarapid nanopore genome sequencing in a critical care setting. New Engl J Med. 2022. doi: 10.1056/NEJMc2112090CrossrefMedlineGoogle Scholar3. Goenka SD, Gorzynski JE, Shafin K, Fisk DG, Pesout T, Monlong J, Jensen TD, Chang P-C, Baid G, Bernstein JA, et al.. Accelerated whole genome nanopore sequencing pipeline enables ultra-rapid identification of disease-causing variants.Nat Biotechnol. 2022. doi: 10.1038/s41587-022-01221-5CrossrefMedlineGoogle Scholar4. Ammirati E, Cipriani M, Moro C, Raineri C, Pini D, Sormani P, Mantovani R, Varrenti M, Pedrotti P, Conca C, et al.; Registro Lombardo delle Miocarditi. Clinical presentation and outcome in a contemporary cohort of patients with acute myocarditis: multicenter lombardy registry.Circulation. 2018; 138:1088–1099. doi: 10.1161/CIRCULATIONAHA.118.035319LinkGoogle Scholar5. Schultheiss HP, Fairweather D, Caforio ALP, Escher F, Hershberger RE, Lipshultz SE, Liu PP, Matsumori A, Mazzanti A, McMurray J, et al.. Dilated cardiomyopathy.Nat Rev Dis Primers. 2019; 5:32. doi: 10.1038/s41572-019-0084-1CrossrefMedlineGoogle Scholar6. Shafin K, Pesout T, Chang PC, Nattestad M, Kolesnikov A, Goel S, Baid G, Kolmogorov M, Eizenga JM, Miga KH, et al.. Haplotype-aware variant calling with PEPPER-Margin-DeepVariant enables high accuracy in nanopore long-reads.Nat Methods. 2021; 18:1322–1332. doi: 10.1038/s41592-021-01299-wCrossrefMedlineGoogle Scholar7. Sedlazeck FJ, Rescheneder P, Smolka M, Fang H, Nattestad M, von Haeseler A, Schatz MC. Accurate detection of complex structural variations using single-molecule sequencing.Nat Methods. 2018; 15:461–468. doi: 10.1038/s41592-018-0001-7CrossrefMedlineGoogle Scholar Previous Back to top Next FiguresReferencesRelatedDetails April 2022Vol 15, Issue 2 Advertisement Article InformationMetrics © 2022 American Heart Association, Inc.https://doi.org/10.1161/CIRCGEN.121.003591PMID: 35133172 Originally publishedFebruary 8, 2022 Keywordsgenomedilatedcardiomyopathydiagnosisgenomicssequence analysis, DNAPDF download Advertisement SubjectsCardiomyopathyGeneticsPrecision Medicine
Whole-genome sequencing (WGS) can identify variants that cause genetic disease, but the time required for sequencing and analysis has been a barrier to its use in acutely ill patients. In the present study, we develop an approach for ultra-rapid nanopore WGS that combines an optimized sample preparation protocol, distributing sequencing over 48 flow cells, near real-time base calling and alignment, accelerated variant calling and fast variant filtration for efficient manual review. Application to two example clinical cases identified a candidate variant in <8 h from sample preparation to variant identification. We show that this framework provides accurate variant calls and efficient prioritization, and accelerates diagnostic clinical genome sequencing twofold compared with previous approaches.
The SARS-CoV-2 pandemic has differentially impacted populations across race and ethnicity. A multi-omic approach represents a powerful tool to examine risk across multi-ancestry genomes. We leverage a pandemic tracking strategy in which we sequence viral and host genomes and transcriptomes from nasopharyngeal swabs of 1049 individuals (736 SARS-CoV-2 positive and 313 SARS-CoV-2 negative) and integrate them with digital phenotypes from electronic health records from a diverse catchment area in Northern California. Genome-wide association disaggregated by admixture mapping reveals novel COVID-19-severity-associated regions containing previously reported markers of neurologic, pulmonary and viral disease susceptibility. Phylodynamic tracking of consensus viral genomes reveals no association with disease severity or inferred ancestry. Summary data from multiomic investigation reveals metagenomic and HLA associations with severe COVID-19. The wealth of data available from residual nasopharyngeal swabs in combination with clinical data abstracted automatically at scale highlights a powerful strategy for pandemic tracking, and reveals distinct epidemiologic, genetic, and biological associations for those at the highest risk.
ABSTRACTThe SARS-CoV-2 pandemic has differentially impacted populations of varied race, ethnicity and socioeconomic status. Admixture mapping and local ancestry inference represent powerful tools to examine genetic risk within multi-ancestry genomes independent of these confounding social constructs. Here, we leverage a pandemic tracking strategy in which we sequence viral and host genomes and transcriptomes from 1,327 nasopharyngeal swab residuals and integrate them with digital phenotypes from electronic health records. We demonstrate over-representation of individuals possessing Oceanian and Indigenous American ancestry in SARS-CoV-2 positive populations. Genome-wide-association disaggregated by admixture mapping reveals regions of chromosomes 5 and 14 associated with COVID19 severity within African and Oceanic local ancestries, respectively, independent of overall ancestry fraction. Phylodynamic tracking of consensus viral genomes reveals no association with disease severity or inferred ancestry. We further present summary data from a multi-omic investigation of human-leukocyte-antigen (HLA) typing, nasopharyngeal microbiome and human transcriptomics that reveal metagenomic and HLA associations with severe COVID19 infection. This work demonstrates the power of multi-omic pandemic tracking and genomic analyses to reveal distinct epidemiologic, genetic and biological associations for those at the highest risk.