The inundating rate of scientific publishing means every researcher will miss new discoveries from overwhelming saturation. To address this limitation, we employ natural language processing to overcome human limitations in reading, curation, and knowledge synthesis, with domain-specific applications to genetics and genomics. We construct a corpus of 3.5 million normalized genetics and genomics abstracts and implement both semantic and network-based embedding models. Our methods not only capture broad biological concepts and relationships but also predict complex phenomena such as gene expression. Through a rigorous temporal validation framework, we demonstrate that our embeddings successfully predict gene-disease associations, cancer driver genes, and experimentally-verified protein interactions years before their formal documentation in literature. Additionally, our embeddings successfully predict experimentally verified gene-gene interactions absent from the literature. These findings demonstrate that substantial undiscovered knowledge exists within the collective scientific literature and that computational approaches can accelerate biological discovery by identifying hidden connections across the fragmented landscape of scientific publishing.
Somatic mutations in individual cells create genomic mosaicism, influencing genetic disorders and cancers. While clonal mutations in cancers are well-studied, rarer somatic variants in normal tissues remain poorly characterized. This study systematically evaluates detection methods using a personalized donor-specific assembly (DSA) from a neurotypical individual's dorsolateral prefrontal cortex assessed with Oxford Nanopore, NovaSeq, linked-read sequencing, Cas9-targeted long-read sequencing (TEnCATS), and single-neuron MALBAC amplification. The haplotype-resolved DSA improved cross-platform analysis, dramatically increasing phasing rates. Germline SNVs, structural variations (SVs), and transposable elements (TEs) were recalled with 99.4%-99.7% accuracy in bulk tissue, and phased haplotype analysis reduced false positives by 15.4%-75.1% for putative somatic candidates. Long-read single-neuron sequencing detected nine somatic SV candidates, demonstrating enhanced sensitivity for rare variants, while TEnCATS identified eight low-frequency somatic TE candidates. These findings highlight advanced methodologies for precise somatic variant detection, critical for understanding mosaicism's role in health and disease.
Somatic mutations in individual cells lead to genomic mosaicism, contributing to the intricate regulatory landscape of genetic disorders and cancers. To evaluate and refine the detection of somatic mosaicism across different technologies with personalized donor-specific assembly (DSA), we obtained tissue from the dorsolateral prefrontal cortex (DLPFC) of a post-mortem neurotypical 31-year-old individual. We sequenced bulk DLPFC tissue using Oxford Nanopore Technologies (∼60X), NovaSeq (∼30X), and linked-read sequencing (∼28X). Additionally, we applied Cas9 capture methodology coupled with long-read sequencing (TEnCATS), targeting active transposable elements. We also isolated and amplified DNA from flow-sorted single DLPFC neurons using MALBAC, sequencing 115 of these MALBAC libraries on Nanopore and 94 on NovaSeq. We constructed a haplotype-resolved assembly with a total length of 5.77 Gb and a phase block length of 2.67 Mb (N50) to facilitate cross-platform analysis of somatic genetic variations. We observed an increase in the phasing rate from 11.6% to 38.0% between short-read and long-read technologies. By generating a catalog of phased germline SNVs, CNVs, and TEs from the assembled genome, we applied standard approaches to recall these variants across sequencing technologies. We achieved aggregated recall rates from 97.3% to 99.4% based on long-read bulk tissue data, setting an upper bound for detection limits. Moreover, utilizing haplotype-based analysis from DSA, we achieved a remarkable reduction in false positive somatic calls in bulk tissue, ranging from 14.9% to 72.4%. We developed pipelines leveraging DSA information to enhance somatic large genetic variant calling in long-read single cells. By examining somatic variation using long-reads in 115 individual neurons, we identified 468 candidate somatic heterozygous large deletions (1.5Mb - 20Mb), 137 of which intersected with short-read single-cell data. Additionally, we identified 61 putative somatic TEs (60 Alu s, one LINE-1) in the single-cell data. Collectively, our analysis spans personalized assembly to single-cell somatic variant calling, providing a comprehensive ab initio ad finem approach and resource in real human tissue.
Merkel cell carcinoma (MCC) is an aggressive disease with poor survival outcomes and increasing incidence. There is a clear and present need for enhanced understanding of cellular mechanisms of tumorigenesis, validation of robust genetic signatures predictive of aggressive disease, and novel informatics tools to simplify analysis of Merkel cell polyomavirus (MCPyV)-host genome interactions. Genomic DNA was harvested from 54 MCC tumors for exome sequencing and in-depth genetic profiling of a 226-gene panel. We further developed a robust informatics package (MCPyViewer) optimized for MCPyV integration site analysis with graphical output to simplify usability for end users. Finally, we assessed the prognostic impact of specific genetic signatures on MCC-specific survival in our cohort. Our study included 54 patients (n = 44 MCPyV positive), 11 (20.4%) of whom had died of MCC at last follow-up. Human genes altered at high frequency included LRP1B (n = 10, 18.5%), FAT1 (n = 9, 16.7%), KMT2D (n = 9, 16.7%), and RB1 (n = 7, 13.0%). In 36 of 44 (81.8%) MCPyV-positive tumors, we identified viral integration into the human genome with a median of two events per tumor. In six tumors, MCPyV integrated into Catalogue of Somatic Mutations in Cancer tier 1 or tier 2 cancer-related human genes. IMPLICATIONS:A combined genomics score incorporating tumor mutational burden and copy-number variation was strongly prognostic of MCC-specific survival controlling for lymph node metastases and tumor MCPyV status; thus, our study adds critical understanding to prognostic markers and tumorigenic mechanisms in MCC.
Survival analyses for prognostic genetic signatures in MCC cohort. A-B. Kaplan-Meier curves for MCC-specific survival by t-TMB category in N+ patients (A) and VP-MCC patients (B). Kaplan-Meier curves for MCC-specific survival by CNV score category in N+ patients (C) and VP-MCC patients (D).
Comparison of t-TMB and CNV scores by clinicopathologic variables in overall cohort. Data shown as median (range) or n (%) for continuous and categorical variables, respectively. Wilcoxon Rank-Sum and Kruskal-Wallis tests used for comparison of categorical and continuous variables, respectively. Additionally, we saw no correlation between primary tumor Breslow depth (mm) and t-TMB (Pearson’s r = -0.131, [95% CI: - 0.412 – 0.173], p = 0.397) or CNV score (Pearson’s r = -0.102, [95% CI: - 0.387 – 0.201], p = 0.511). a Data for primary tumor LVI and Breslow depth available for 45 of 54 (83.3%) patients.
Relative read depth across the MCPyV genome in each VP-MCC tumor. Left column shows MCPyV genome position (RefSeq NC_010277).
Mobile element insertions (MEI) shape the human genome in both germline and somatic tissues. While inherited MEIs are well characterized, mapping somatic MEIs (sMEI) in non-cancer tissues remains challenging due to their low allelic fraction and repetitive nature. We established an integrative framework for sMEI analysis leveraging modern sequencing technologies and analytical innovations. We first benchmarked sMEI detection and demonstrated advantages of long-read and MEI-targeted sequencing for ultra-low-frequency events using a mixture of well-established cell lines. We then showed that haplotype phasing and donor-specific assemblies refine sMEI detection, effectively distinguishing from germline and false signals in in-silico tumor-normal mixtures. We further developed a source-tracing strategy based on internal sequence variation, expanding the catalogue of active source elements beyond traditional transduction-based methods. Applying this framework to donor tissues, we identified 18 rare somatic L1 insertions, revealing structural and source diversity. Our work provides a foundational framework and biological insight into sMEIs.
Summary of definitive treatment modalities for MCC patients in our cohort, stratified by stage. Data presented as n (%). a Including sentinel lymph node biopsy ± completion lymphadenectomy of regional lymphatic basin.
Representative MCPyV integration events in MiOTO_4055 and MiOTO_4003 that fell within human oncogenes. Red arrow: MCPyV integration; Black segment: contigs assembled at integration events; Segments with black arrowhead: zoomed in contigs aligned to the human genome; Segments with red arrowhead: zoomed in contigs aligned to the MCPyV genome. Direction of segments indicate the strand of contigs aligned to. Note figure constructed using MCPyViewer.
The transfer of mitochondrial DNA into the nuclear genomes of eukaryotes (Numts) has been linked to lifespan in nonhuman species and recently demonstrated to occur in rare instances from one human generation to the next. Here, we investigated numtogenesis dynamics in humans in 2 ways. First, we quantified Numts in 1,187 postmortem brain and blood samples from different individuals. Compared to circulating immune cells (n = 389), postmitotic brain tissue (n = 798) contained more Numts, consistent with their potential somatic accumulation. Within brain samples, we observed a 5.5-fold enrichment of somatic Numt insertions in the dorsolateral prefrontal cortex (DLPFC) compared to cerebellum samples, suggesting that brain Numts arose spontaneously during development or across the lifespan. Moreover, an increase in the number of brain Numts was linked to earlier mortality. The brains of individuals with no cognitive impairment (NCI) who died at younger ages carried approximately 2 more Numts per decade of life lost than those who lived longer. Second, we tested the dynamic transfer of Numts using a repeated-measures whole-genome sequencing design in a human fibroblast model that recapitulates several molecular hallmarks of aging. These longitudinal experiments revealed a gradual accumulation of 1 Numt every ~13 days. Numtogenesis was independent of large-scale genomic instability and unlikely driven by cell clonality. Targeted pharmacological perturbations including chronic glucocorticoid signaling or impairing mitochondrial oxidative phosphorylation (OxPhos) only modestly increased the rate of numtogenesis, whereas patient-derived SURF1-mutant cells exhibiting mtDNA instability accumulated Numts 4.7-fold faster than healthy donors. Combined, our data document spontaneous numtogenesis in human cells and demonstrate an association between brain cortical somatic Numts and human lifespan. These findings open the possibility that mito-nuclear horizontal gene transfer among human postmitotic tissues produces functionally relevant human Numts over timescales shorter than previously assumed.
Diverse sets of complete human genomes are required to construct a pangenome reference and to understand the extent of complex structural variation. Here, we sequence 65 diverse human genomes and build 130 haplotype-resolved assemblies (130 Mbp median continuity), closing 92% of all previous assembly gaps1,2 and reaching telomere-to-telomere (T2T) status for 39% of the chromosomes. We highlight complete sequence continuity of complex loci, including the major histocompatibility complex (MHC), SMN1/SMN2, NBPF8, and AMY1/AMY2, and fully resolve 1,852 complex structural variants (SVs). In addition, we completely assemble and validate 1,246 human centromeres. We find up to 30-fold variation in α-satellite high-order repeat (HOR) array length and characterize the pattern of mobile element insertions into α-satellite HOR arrays. While most centromeres predict a single site of kinetochore attachment, epigenetic analysis suggests the presence of two hypomethylated regions for 7% of centromeres. Combining our data with the draft pangenome reference1 significantly enhances genotyping accuracy from short-read data, enabling whole-genome inference3 to a median quality value (QV) of 45. Using this approach, 26,115 SVs per sample are detected, substantially increasing the number of SVs now amenable to downstream disease association studies.
When somatic cells acquire complex karyotypes, they often are removed by the immune system. Mutant somatic cells that evade immune surveillance can lead to cancer. Neurons with complex karyotypes arise during neurotypical brain development, but neurons are almost never the origin of brain cancers. Instead, somatic mutations in neurons can bring about neurodevelopmental disorders, and contribute to the polygenic landscape of neuropsychiatric and neurodegenerative disease. A subset of human neurons harbors idiosyncratic copy number variants (CNVs, "CNV neurons"), but previous analyses of CNV neurons are limited by relatively small sample sizes. Here, we develop an allele-based validation approach, SCOVAL, to corroborate or reject read-depth based CNV calls in single human neurons. We apply this approach to 2,125 frontal cortical neurons from a neurotypical human brain. SCOVAL identifies 226 CNV neurons, which include a subclass of 65 CNV neurons with highly aberrant karyotypes containing whole or substantial losses on multiple chromosomes. Moreover, we find that CNV location appears to be nonrandom. Recurrent regions of neuronal genome rearrangement contain fewer, but longer, genes.
Somatic mosaicism is defined as an occurrence of two or more populations of cells having genomic sequences differing at given loci in an individual who is derived from a single zygote. It is a characteristic of multicellular organisms that plays a crucial role in normal development and disease. To study the nature and extent of somatic mosaicism in autism spectrum disorder, bipolar disorder, focal cortical dysplasia, schizophrenia, and Tourette syndrome, a multi-institutional consortium called the Brain Somatic Mosaicism Network (BSMN) was formed through the National Institute of Mental Health (NIMH). In addition to genomic data of affected and neurotypical brains, the BSMN also developed and validated a best practices somatic single nucleotide variant calling workflow through the analysis of reference brain tissue. These resources, which include >400 terabytes of data from 1087 subjects, are now available to the research community via the NIMH Data Archive (NDA) and are described here.
When somatic cells acquire complex karyotypes, they are removed by the immune system. Mutant somatic cells that evade immune surveillance can lead to cancer. Neurons with complex karyotypes arise during neurotypical brain development, but neurons are almost never the origin of brain cancers. Instead, somatic mutations in neurons can bring about neurodevelopmental disorders, and contribute to the polygenic landscape of neuropsychiatric and neurodegenerative disease. A subset of human neurons harbors idiosyncratic copy number variants (CNVs, "CNV neurons"), but previous analyses of CNV neurons have been limited by relatively small sample sizes. Here, we developed an allele-based validation approach, SCOVAL, to corroborate or reject read-depth based CNV calls in single human neurons. We applied this approach to 2,125 frontal cortical neurons from a neurotypical human brain. This approach identified 226 CNV neurons, as well as a class of CNV neurons with complex karyotypes containing whole or substantial losses on multiple chromosomes. Moreover, we found that CNV location appears to be nonrandom. Recurrent regions of neuronal genome rearrangement contained fewer, but longer, genes.
Detailed summary of host genome somatic INDELs identified by targeted capture sequencing analysis of FFPE tumor specimens