Somatic mutations that arise post-zygotically create genetic diversity among normal human cells and provide key insights into human development and aging. Fibroblast-derived induced pluripotent stem cells (iPSCs) have proved to be a useful system for disease modelling; however, due to their clonal nature, iPSC lines carry somatic mutations inherited from the founder cells, raising concerns about their genomic integrity. At the same time, this clonality enables single-cell-level discovery of somatic mutations and the reconstruction of developmental lineages. In living individuals, though, this approach requires invasive biopsies and is limited to skin-derived lineages. Here, we generated 33 urine-derived iPSC lines from four males representing two father- son relationships, performed shallow whole-genome sequencing of the lines and analyzed somatic mutations. Derived iPSCs representing single cells from urine carried a few hundred of somatic single-nucleotide variants per genome, dominated by endogenous, clock-like mutational signatures and lacking environmental imprints such as UV-associated mutations. Copy-number analysis identified somatic CNVs in most of the lines and revealed higher CNV burdens in fathers than in sons, consistent with age-related structural mosaicism. Shared mutations across lines enabled reconstruction of cell lineage phylogenetic trees. In summary, urine-derived iPSCs showed genomic alterations comparable to those in fibroblast-derived iPSC lines and represent a valuable non-invasive alternative for disease modeling. Overall, this study provides the first genome-wide characterization of somatic mutations in urine-derived iPSCs and establishes them as a practical and non-invasive platform for charting somatic mutation landscapes and tracing developmental lineages in living humans.
Abstract Background: Most colorectal cancer (CRC) arises from polyps and is mainly prevented by polypectomy. The most important polyps to manage with colonoscopy are those with highest CRC risk- namely, advanced polyps (> 1cm, villous histology or high-grade dysplasia (HGD). Yet, 48% of advanced polyps recur within 1 to 3 years of removal, and up to 5% of advanced polyps under surveillance still progress to CRC. We performed this pilot study to investigate molecular and microenvironment heterogeneities of polyps and how they might impact their clinical behavior. Methods: GeoMx spatial transcriptomics was performed on FFPE tissues from three different polyp outcome phenotypes (POPs) including the polyp that does not recur (POP-NR), that recurs following polypectomy but cured by colonoscopy (POP-R) or the polyp despite polypectomy develops CRC at the polypectomy(ies) site (POP-CRC). Normal colon and polyp with low or HGD from the index polyp were assessed from 6 patients with POP-NR, 9 with POP-R and 12 with POP-CRC with a minimum of 2 follow up colonoscopies at 3-year intervals. Epithelium was identified as PanCK positive, and stroma identified as PanCK negative and positive nuclear staining. Cell type deconvolution was implemented using single-cell adult human intestinal tract catalogue (Elmentaite et al., https://www.gutcellatlas.org/). Differential gene expression (DEG), cell type abundance and functional enrichment analyses was performed using linear mixed effects model, and observations with p-value < 0.05 and log2FC >= |1| are reported. Results: The greatest number of DEGs were identified in stroma of the index POP-CRC compared to POP-NR polyps (n = 17 down and 11 upregulated genes), followed by epithelium of the index POP-CRC vs POP-NR polyps (n = 10 down and one upregulated gene(s)). Between the index POP-CRC and POP-R, epithelium showed downregulation of 7 genes, and upregulation of no genes. In stroma two genes were down and none upregulated. Four genes were downregulated in stroma of the index POP-R vs POP-NR polyp and 3 in epithelium, and one gene was upregulated in stroma and another in epithelium. MZT2B was upregulated in stroma of both the index POP-R (log2FC = 1.09, p-value = 0.0014) and POP-CRC (log2FC = 1.29, p-value = 0.0001) and has been implicated as a prognostic marker associated with worse prognosis in certain cancers. IgM and IgA plasma cells, proximal progenitors, MMP9+ inflammatory macrophages and myofibroblasts were found to be downregulated in epithelium of index POP-CRC and POP-R compared to POP-NR polyps; while microfold cells were downregulated in stroma of index POP-CRC and POP-R compared to POP-NR polyps. Conclusions: The stromal microenvironment and less so, epithelial features present in the index polyp differ based on a polyp’s future clinical behavior. The polyp-immune interaction warrants further study as a potential prevention target against polyp progression. Citation Format: Mrunal Dehankar, Lisa A. Boardman, Daniel O'Brien, E. Aubrey Thompson, Jennifer M. Kachergus, Ji Shi, Alexej Abyzov, Milovan Suvakov, Rondell P. Graham, Chen Wang. Stromal immune composition significantly contributes to adenomatous polyp recurrence or progression to cancer [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 4018.
BACKGROUND:Current literature suggestsisocitrate dehydrogenase (IDH)-mutant astrocytoma contains several molecular subgroups. In this study, we are interested in determining the connection between different molecular subgroups with grade and/or survival. METHODS:A cohort of 470 Mayo Clinic adult patients (≥18 years, 56.2% male) with primary IDH-mutant astrocytoma diagnosed by World Health Organization (WHO) 2021 criteria were examined. Results were validated in an independent cohort of 614 Mayo Clinic Neuropathology consult patients and 235 The Cancer Genome Atlas (TCGA) patients. RESULTS:The Mayo Clinic Practice cohort confirmed the association of CDKN2A/B deletion with overall survival (OS, homozygous vs hemizygous vs intact, 2.7 vs 9.6 vs 17.2 years, P < .001). Phosphatase and tensin homolog (PTEN) deletion was also associated with poor OS (7.3 vs 17.4 years, P < .001). Increased number of copy number alterations was associated with OS (continuous variable, HR = 1.027, P < .001). Carrying one or more copies of the germline risk allele at rs55705857 was associated with earlier age of onset (median age 33 vs 35 years, P = .01), and a shorter OS after adjusting for age, grade, sex and treatment (HR = 1.81, P = .007). The Mayo Clinic Neuropathology Consult cohort and TCGA were utilized to validate age of onset and survival, respectively. Unsupervised clustering of the copy number alterations identified several clinically significant groups that may define pathways to disease progression. Losses of chromosomes 11p, 13q, 1p, and 10q were all associated with reduced overall survival in the Mayo Clinic cohort. CONCLUSIONS:Patients with hemizygous loss of CDKN2A/B, loss of PTEN, increased number of copy number alterations, specific chromosomal arm losses or rs55705857 germline risk allele have reduced overall survival.
Cell differentiation involves shifts in chromatin organization allowing transcription factors (TFs) to bind enhancer elements and modulate gene expression. The TF-enhancer-gene regulatory interactions that control the formation of neuronal lineages have yet to be charted in humans. Here, we mapped enhancer elements and conducted an integrative analysis of epigenomic and transcriptomic profiles across 60 days of differentiation of human forebrain organoids derived from 10 individuals with autism spectrum disorder (ASD) and their neurotypical fathers. This multi-omics profiling allowed us to build an enhancer-driven gene regulatory network (GRN) of early neural development. We validated the GRN by performing a loss-of-function experiment with FOXG1 – one of the master TFs in the development of the mammalian brain. Analysis of the constructed GRN identified regulatory hierarchies driving the specification of neuronal cell types. By analyzing differential gene expression in ASD through the GRN hierarchy, we associated the ASD transcriptomic signatures to altered activity of key TFs. We found that macrocephalic ASD was principally driven by an increased activity of BHLHE22, FOXG1, EOMES, and NEUROD2, which are major regulators of excitatory neuron fate. Normocephalic ASD, on the contrary, was driven by decreased activity of those same TFs and by an increased activity of LMX1B and FOXB1 – two upstream TF repressors of FOXG1. These findings suggest that ASD is characterized by an altered early gene regulatory program that specifies neuronal cell lineages in the fetal brain. Thus, constructing a GRN of early brain development modeled in organoids provides insights into the etiology of ASD and can guide future experimental approaches to establish its genetic causes and treatment strategies.
Xeroderma Pigmentosum group C (XP-C) is an autosomal recessive disorder caused by mutations in the XPC gene, leading to defective nucleotide excision repair. This defect leads to genomic instability and a profound cancer predisposition. To model this disease, we generated induced pluripotent stem cells (iPSCs) from an XP-C patient carrying a novel homozygous nonsense mutation in the XPC gene (c.1830C>A). The resulting iPSCs demonstrated typical pluripotent characteristics, including expression of key markers and trilineage differentiation capability. However, genomic assessment revealed progressive karyotypic instability during extended culture. While initial whole-genome sequencing detected no major chromosomal abnormalities, subsequent G-banding analysis identified acquired trisomy 12 in two lines (CL12 and CL27) and a derivative X chromosome in a third line (CL30). These abnormalities were absent in early-passage analyses, indicating that they were acquired and selected for during extended culture. The acquisition of a derivative X chromosome in CL30, alongside recurrent trisomy 12, represents a unique cytogenetic signature likely attributable to the underlying XPC defect. We hypothesize that the loss of GG-NER creates a permissive genomic environment, accelerating the accumulation of DNA damage and chromosomal missegregation under replicative stress. This temporal divergence in genomic integrity highlights how culture pressures drive chromosomal evolution in XP-C iPSCs independently of initial reprogramming. Our findings emphasize that XP-C iPSCs require continuous genomic surveillance and provide a model for investigating how DNA repair deficiencies interact with in vitro culture stress.
The repertoire of neurons and their progenitors depends on their location along the antero-posterior and dorso-ventral axes of the neural tube. To model these axes, we designed the Dual Orthogonal-Morphogen Assisted Patterning System (Duo-MAPS) diffusion device to expose spheres of induced pluripotent stem cells (iPSCs) to concomitant orthogonal gradients of a posteriorizing and a ventralizing morphogen, activating WNT and SHH signaling, respectively. Comparison with single-cell transcriptomes from the fetal human brain revealed that Duo-MAPS-patterned organoids generated an extensive diversity of neuronal lineages from the forebrain, midbrain, and hindbrain. WNT and SHH crosstalk translated into early patterns of gene expression programs associated with the generation of specific brain lineages with distinct functional networks. Human iPSC lines showed substantial interindividual and line-to-line variations in their response to morphogens, highlighting that genetic and epigenetic variations may influence regional specification. Morphogen gradients promise to be a key approach to model the brain in its entirety.
Tourette syndrome (TS) is a disorder of high-order integration of sensory, motor, and cognitive functions afflicting as many as 1 in 150 children and characterized by motor hyperactivity and tics. Despite high familial recurrence rates, a few risk genes and no biomarkers have emerged as causative or predisposing factors. The syndrome is believed to originate in basal ganglia, where patterns of motor programs are encoded. Postmortem immunocytochemical analyses of brains with severe TS revealed decreases in cholinergic, fast-spiking parvalbumin, and somatostatin interneurons within the striatum (caudate and putamen nuclei). Here, we performed single cell transcriptomic and chromatin accessibility analyses of the caudate nucleus from 6 adult TS and 6 control post-mortem brains. The data reproduced the known cellular composition of the adult human striatum, including a majority of medium spiny neurons (MSN) and small populations of GABAergic and cholinergic interneurons. Comparative analysis revealed that interneurons were decreased by roughly 50% in TS brains, while no difference was observed for other cell types. Differential gene expression analysis suggested that mitochondrial function, and specifically oxidative metabolism, in MSN and synaptic function in interneurons are both impaired in TS subjects. Furthermore, such an impairment was coupled with activation of immune response pathways in microglia. Also, our data explicitly link gene expression changes to changes in cis-regulatory activity in the corresponding cell types, suggesting de-regulation as a factor for the etiology of TS. These findings expand on previous research and suggest that impaired modulation of striatal function by interneurons may be the origin of TS symptoms.
Somatic mosaicism is increasingly recognized as a fundamental feature of human biology, yet the detection of somatic mutations remains challenging. The SMaHT Network conducted four large-scale benchmarking experiments to evaluate sequencing technologies, experimental approaches, and computational methods for detecting diverse somatic mutations. Cumulative sequencing coverage exceeded 1,000× with short reads and 100-400× with long reads for each of nine analyzed samples. We defined optimal strategies for integrating bulk short- and long-read sequencing for mutation detection and demonstrated that using donor-specific assemblies and human pangenome improved variant calling and extended mutation catalogs to challenging genomic regions. We benchmarked six duplex-seq technologies and showed that single-cell sequencing resolves cell type-specific mutational patterns and heterogeneity. Our results indicate that bulk, single-cell, and duplex analyses are complementary - and leveraging all three provides comprehensive characterization of mosaicism within a tissue. Together, these findings provide a roadmap for accurate, genome-wide somatic mutation discovery and analysis.
The role of somatic mutations in human development and disease is obscured by difficulties in characterizing mutations at the single cell level and identifying cell types carrying them. Here we analysed somatic genomes of clonal iPSC lines and of single-cells after whole-genome amplification (scWGA) by PTA and ResolveOme from skin fibroblasts, blood and urine of a live donor. Mutation burden and spectra converged across approaches, revealing heterogeneous mutational footprints across cells driven by environmental exposures (UV damage and chemotherapy) and lymphocyte differentiation. Aneuploidies in single cells were detected by all the approaches and were orthogonally validated by Strand-seq. Uniquely, ResolveOme enabled cell-type identification using single-cell transcriptomes. Using a newly developed method accounting for noise and allele drop-out in scWGA, we de novo reconstructed the cell phylogenetic tree for this donor. Together, scWGA establishes a powerful foundation for comprehensive, cell type-aware, lineage-aware profiling of somatic mutations at single cell level.
The WHO 2021 classification of central nervous system tumors incorporated molecular markers to better define diffuse glioma subgroups. We analyzed array-based copy number, targeted next-generation sequencing, germline genotyping, and Illumina EPIC array methylation data from 404 primary IDH-mutant astrocytomas to further understand the relationship between copy number and outcomes. Loss of chromosome arm 13q was associated with poor overall survival after adjusting for age, sex, tumor grade, and treatment status (hazard ratio [HR]=2.04 (95% confidence interval: 1.25-3.33), p=0.004). Chromosome 13q loss was also significantly associated with the total number of copy number alterations (p<0.001). Loss of arm 11p was associated with poor survival in the adjusted analysis (HR=2.29 (1.38-3.79), p=0.001). Unsupervised clustering of the raw copy number data identified 12 clusters, where each cluster was defined by distinct copy number alteration(s). Clusters that contained 13q loss exhibited worse overall survival compared to other clusters. The clusters were significantly associated with age at diagnosis (p=0.015), tumor grade (p<0.001) and copy number complexity (p<0.001). The clusters with a high proportion of grade 4 tumors and with the highest complexity had the poorest overall survival. These findings highlight the prognostic relevance of chromosome 13q and 11p losses in IDH-mutant astrocytomas and support the integration of additional copy number alterations into molecular classification frameworks of IDH-mutant astrocytomas to improve risk stratification and patient management.
Gene conversion is a specific form of homologous recombination (HR), involving the unidirectional transfer of genetic information from one genomic locus to another. CRISPR-Cas9-directed double strand breaks (DSBs) induce both interallelic and interlocus gene conversion in early human embryos and somatic cells, suggesting its potential for correcting pathogenic mutations. However, the key features in mitotic gene conversion, including its efficiency, the length of conversion track, and its dependency on specific recombination proteins, remain largely undefined. Here, we show that allele-specific CRISPR-Cas9-induced DSBs, without exogenous donor templates, can efficiently correct a heterozygous pathogenic variant (c.1582C>T; p.Arg528Trp) in the ATAD3A gene of patient-derived induced pluripotent stem cells (iPSCs). Amplicon-based next-generation sequencing (NGS) revealed that approximately 38%~53% of edited iPSCs carried two wild-type ATAD3A alleles. Notably, over 99% of the corrected alleles derived from the homologous chromosome, indicating that the repair occurred mainly via interallelic gene conversion. Long-range amplicon nanopore sequencing coupled with haplotype analysis showed that the majority of gene conversion tracts was less than 2 kilobases in length. Whole-genome sequencing of three corrected iPSC clones showed the absence of large deletions or structural rearrangements at the ATAD3A target site. However, one clone carried a heterozygous deletion in ATAD3B locus, suggesting that CRISPR-Cas9 can introduce off-target genomic alterations. Knockdown of key HR proteins, including RAD51, CtIP, and BRCA1/2, significantly reduced the correction efficiency, indicating that the gene conversion relies on a RAD51-dependent HR pathway. Together, our findings provide compelling evidence that template-free CRISPR-Cas9-mediated interallelic gene conversion can be harnessed to correct disease causing variants in human iPSCs.
The repertory of neurons generated by progenitor cells depends on their location along antero-posterior and dorso-ventral axes of the neural tube. To understand if recreating those axes was sufficient to specify human brain neuronal diversity, we designed a mesofluidic device termed Duo-MAPS to expose induced pluripotent stem cells (iPSC) to concomitant orthogonal gradients of a posteriorizing and a ventralizing morphogen, activating WNT and SHH signaling, respectively. Comparison of single cell transcriptomes with fetal human brain revealed that Duo-MAPS-patterned organoids generated the major neuronal lineages of the forebrain, midbrain, and hindbrain. Morphogens crosstalk translated into early patterns of gene expression programs predicting the generation of specific brain lineages. Human iPSC lines from six different genetic backgrounds showed substantial differences in response to morphogens, suggesting that interindividual genomic and epigenomic variations could impact brain lineages formation. Morphogen gradients promise to be a key approach to model the brain in its entirety.
SUMMARY:Copy number variation (CNV) and alteration (CNA) analysis is a crucial component in many genomic studies and its applications span from basic research to clinic diagnostics and personalized medicine. CNVpytor is a tool featuring a read depth-based caller and combined read depth and B-allele frequency (BAF) based 2D caller to find CNVs and CNAs. The tool stores processed intermediate data and CNV/CNA calls in a compact HDF5 file-pytor file. Here, we describe a new track in igv.js that utilizes pytor and whole genome variant files as input for on-the-fly read depth and BAF visualization, CNV/CNA calling and analysis. Embedding into HTML pages and Jupiter Notebooks enables convenient remote data access and visualization simplifying interpretation and analysis of omics data. AVAILABILITY AND IMPLEMENTATION:The CNVpytor track is integrated with igv.js and available at https://github.com/igvteam/igv.js. The documentation is available at https://github.com/igvteam/igv.js/wiki/cnvpytor. Usage can be tested in the IGV-Web app at https://igv.org/app and also on https://github.com/abyzovlab/CNVpytor.
We developed a generally applicable method, CRISPR/Cas9-targeted long-read sequencing (CTLR-Seq), to resolve, haplotype-specifically, the large and complex regions in the human genome that had been previously impenetrable to sequencing analysis, such as large segmental duplications (SegDups) and their associated genome rearrangements. CTLR-Seq combines in vitro Cas9-mediated cutting of the genome and pulse-field gel electrophoresis to isolate intact large (i.e., up to 2,000 kb) genomic regions that encompass previously unresolvable genomic sequences. These targets are then sequenced (amplification-free) at high on-target coverage using long-read sequencing, allowing for their complete sequence assembly. We applied CTLR-Seq to the SegDup-mediated rearrangements that constitute the boundaries of, and give rise to, the 22q11.2 Deletion Syndrome (22q11DS), the most common human microdeletion disorder. We then performed de novo assembly to resolve, at base-pair resolution, the full sequence rearrangements and exact chromosomal breakpoints of 22q11.2DS (including all common subtypes). Across multiple patients, we found a high degree of variability for both the rearranged SegDup sequences and the exact chromosomal breakpoint locations, which coincide with various transposons within the 22q11.2 SegDups, suggesting that 22q11DS can be driven by transposon-mediated genome recombination. Guided by CTLR-Seq results from two 22q11DS patients, we performed three-dimensional chromosomal folding analysis for the 22q11.2 SegDups from patient-derived neurons and astrocytes and found chromosome interactions anchored within the SegDups to be both cell type-specific and patient-specific. Lastly, we demonstrated that CTLR-Seq enables cell-type specific analysis of DNA methylation patterns within the deletion haplotype of 22q11DS.
Regulation of gene expression through enhancers is one of the major processes shaping the structure and function of the human brain during development. High-throughput assays have predicted thousands of enhancers involved in neurodevelopment, and confirming their activity through orthogonal functional assays is crucial. Here, we utilized Massively Parallel Reporter Assays (MPRAs) in stem cells and forebrain organoids to evaluate the activity of ~7,000 gene-linked enhancers previously identified in human fetal tissues and brain organoids. We used a Gaussian mixture model to evaluate the contribution of background noise in the measured activity signal to confirm the activity of ~35% of the tested enhancers, with most showing temporal-specific activity, suggesting their evolving role in neurodevelopment. The temporal specificity was further supported by the correlation of activity with gene expression. Our findings provide a valuable gene regulatory resource to the scientific community.
Little is known about the origin of germ cells in humans. We previously leveraged post-zygotic mutations to reconstruct zygote-rooted cell lineage ancestry trees in a phenotypically normal woman, termed NC0. Here, by sequencing the genome of her children and their father, we analyze the transmission of early pre-gastrulation lineages and corresponding mutations across human generations. We find that the germline in NC0 is polyclonal and is founded by at least two cells likely descending from the two blastomeres arising from the first zygotic cleavage. Analyzes of public data from several multi-children families and from 1934 familial quads confirm this finding in larger cohorts, revealing that known imbalances of up to 90:10 in early lineages allocation in somatic tissues are not reflected in mutation transmission to offspring, establishing a fundamental difference in lineage allocation between the soma and the germline. Analyzes of all the data consistently suggest that the germline has a balanced 50:50 lineage allocation from the first two blastomeres. The origin of germ cells in humans remains elusive. Here, the authors use post-zygotic mutations to trace cell lineages across human generations, finding a fundamental difference between soma and germline in lineage allocation and suggesting a 50:50 contribution of the first two blastomeres to the germline.
When somatic cells acquire complex karyotypes, they often are removed by the immune system. Mutant somatic cells that evade immune surveillance can lead to cancer. Neurons with complex karyotypes arise during neurotypical brain development, but neurons are almost never the origin of brain cancers. Instead, somatic mutations in neurons can bring about neurodevelopmental disorders, and contribute to the polygenic landscape of neuropsychiatric and neurodegenerative disease. A subset of human neurons harbors idiosyncratic copy number variants (CNVs, "CNV neurons"), but previous analyses of CNV neurons are limited by relatively small sample sizes. Here, we develop an allele-based validation approach, SCOVAL, to corroborate or reject read-depth based CNV calls in single human neurons. We apply this approach to 2,125 frontal cortical neurons from a neurotypical human brain. SCOVAL identifies 226 CNV neurons, which include a subclass of 65 CNV neurons with highly aberrant karyotypes containing whole or substantial losses on multiple chromosomes. Moreover, we find that CNV location appears to be nonrandom. Recurrent regions of neuronal genome rearrangement contain fewer, but longer, genes.
Germline mutations modulate the risk of developing schizophrenia (SCZ). Much less is known about the role of mosaic somatic mutations in the context of SCZ. Deep (239×) whole-genome sequencing (WGS) of brain neurons from 61 SCZ cases and 25 controls postmortem identified mutations occurring during prenatal neurogenesis. SCZ cases showed increased somatic variants in open chromatin, with increased mosaic CpG transversions (CpG>GpG) and T>G mutations at transcription factor binding sites (TFBSs) overlapping open chromatin, a result not seen in controls. Some of these variants alter gene expression, including SCZ risk genes and genes involved in neurodevelopment. Although these mutational processes can reflect a difference in factors indirectly involved in disease, increased somatic mutations at developmental TFBSs could also potentially contribute to SCZ.
Somatic mosaicism is defined as an occurrence of two or more populations of cells having genomic sequences differing at given loci in an individual who is derived from a single zygote. It is a characteristic of multicellular organisms that plays a crucial role in normal development and disease. To study the nature and extent of somatic mosaicism in autism spectrum disorder, bipolar disorder, focal cortical dysplasia, schizophrenia, and Tourette syndrome, a multi-institutional consortium called the Brain Somatic Mosaicism Network (BSMN) was formed through the National Institute of Mental Health (NIMH). In addition to genomic data of affected and neurotypical brains, the BSMN also developed and validated a best practices somatic single nucleotide variant calling workflow through the analysis of reference brain tissue. These resources, which include >400 terabytes of data from 1087 subjects, are now available to the research community via the NIMH Data Archive (NDA) and are described here.
Mosaic variants (MVs) reflect mutagenic processes during embryonic development and environmental exposure, accumulate with aging and underlie diseases such as cancer and autism. The detection of noncancer MVs has been computationally challenging due to the sparse representation of nonclonally expanded MVs. Here we present DeepMosaic, combining an image-based visualization module for single nucleotide MVs and a convolutional neural network-based classification module for control-independent MV detection. DeepMosaic was trained on 180,000 simulated or experimentally assessed MVs, and was benchmarked on 619,740 simulated MVs and 530 independent biologically tested MVs from 16 genomes and 181 exomes. DeepMosaic achieved higher accuracy compared with existing methods on biological data, with a sensitivity of 0.78, specificity of 0.83 and positive predictive value of 0.96 on noncancer whole-genome sequencing data, as well as doubling the validation rate over previous best-practice methods on noncancer whole-exome sequencing data (0.43 versus 0.18). DeepMosaic represents an accurate MV classifier for noncancer samples that can be implemented as an alternative or complement to existing methods.