Purifying selection of mtDNA mutations is a vital process that cleanses the mitochondrial genome of detrimental variants that may endanger individuals and populations. A common measure of purifying selection is the increase of average synonymity by reduction the proportion of mostly detrimental non-synonymous mutations. The mechanisms underlying purifying selection are still debated. The Makova group has recently published high- fidelity analysis of mtDNA mutations in individual human oocytes (Arbeithuber et al., 2025). The authors observed a decrease in the proportion of potentially detrimental coding and conservative mutations at higher mutant fractions (MFs) and interpreted this as purifying selection removing detrimental mutations at higher MFs. We noted, however, that, in contrast to what would be expected under purifying selection, the synonymity of oocyte mutations was very low and decreased, rather than increased, at higher MFs. We hypothesized that this inconsistency resulted from non-synonymous mutations being prone to strong positive selection which erroneously made coding mutations appear negatively selected in comparison. In support of our hypothesis, we show that non-coding oocytes mutations indeed are under strong positive selection. To alleviate this setback, we reanalyzed the data using a new metric of intracellular clonal selection and neutral synonymous mutations as the reference. We demonstrated that coding mutations are in fact under prevailing positive selection. This is in line with previous estimates of positive selection in primordial germ cells (PGCs) and in mother-child pairs. Importantly, prevailing positive selection does not imply the absence of negative selection. We show that specific types of mutations may be under prevailing purifying selection (e.g., the Co1 gene). Of note, this prevailing positive selection pertains only to the most recent, germline mtDNA mutations which have not been yet inherited into the next generation. Purifying selection steps in as germline mutations proceed to subsequent generations. The implications of these findings and the potential benefits of positive selection of detrimental mtDNA mutations are discussed. ### Competing Interest Statement The authors have declared no competing interest.
Somatic mitochondrial DNA (mtDNA) mutations are frequently observed in tumors, yet their role in pediatric cancers remains poorly understood. The heteroplasmic nature of mtDNA-where mutant and wild-type mtDNA coexist-complicates efforts to define its contribution to disease progression. In this study, bulk whole-genome sequencing of 637 matched tumor-normal samples from the Pediatric Cancer Genome Project revealed an enrichment of functionally impactful mtDNA variants in specific pediatric leukemia subtypes. Collectively, the results from single-cell sequencing of five diagnostic leukemia samples demonstrated that somatic mtDNA mutations can arise early in leukemogenesis and undergo positive selection during disease progression, achieving intermediate heteroplasmy-a "sweet spot" that balances mitochondrial dysfunction with cellular fitness. Network-based systems biology analyses link specific heteroplasmic mtDNA mutations to metabolic reprogramming and therapy resistance. We reveal somatic mtDNA mutations as a potential source of functional heterogeneity and cellular diversity among leukemic cells, influencing their fitness and shaping disease progression.
Mitochondrial DNA (mtDNA) mutagenesis remains poorly understood despite its crucial role in disease, aging, and evolutionary tracing. In this study, we reconstructed a comprehensive 192-component mtDNA mutational spectrum for chordates by analyzing 118,397 synonymous mutations in the CytB gene across 1,697 species and five classes. This analysis revealed three primary forces shaping mtDNA mutagenesis: (i) symmetrical, replication-driven errors by mitochondrial polymerase (POLG), resulting in C > T and A > G mutations that are highly conserved across classes; (ii) asymmetrical, damage-driven C > T mutations on the single-stranded heavy strand with clock-like dynamics; and (iii) asymmetrical A > G mutations on the heavy strand, with dynamics suggesting sensitivity to oxidative damage. The third component, sensitive to oxidative damage, positions mtDNA mutagenesis as a promising marker for metabolic and physiological processes across various classes, species, organisms, tissues, and cells. The deconvolution of the mutational spectra into mutational signatures uncovered deficiencies in both base excision repair (BER) and mismatch repair (MMR) pathways. Further analysis of mutation hotspots, abasic sites, and mutational asymmetries underscores the critical role of single-stranded DNA damage (components ii and iii), which, uncorrected due to BER and MMR deficiencies, contributes roughly as many mutations as POLG-induced errors (component i).
The extent to which somatic mitochondrial DNA (mtDNA) mutations are subject to selection is a fundamental question relevant to development, mitochondrial disease, cancer, and aging. Recently a study from the Sudmant laboratory that used an advanced, high fidelity mutational analysis reported that somatic mutations in protein-coding genes exhibit signatures of negative selection. This report came as surprise as several other studies including those that used same technology reported either lack of selection or positive (destructive) selection on somatic mutations. We hypothesized that these discrepancies may stem, in part, from the inclusion of germline mutations in addition to somatic ones, which could bias selection analyses due to the high synonymity of the latter. To test this, we reanalyzed the Sudmant dataset by separating mutations into germline (defined as shared between related animals) and somatic (not shared between tissues of an animal). We then employed a cumulative curve approach to assess selection without bias. Our analysis reveals that, indeed, an apparent purifying selection signal is driven by an admixture of synonymous germline mutations and disappears upon their removal. The remaining somatic mutations for most part show overall dynamics consistent with neutral drift. However, mutations at higher mutant fractions show positive selection trend, most compatible with a low proportion of mutations experiencing positive selection. While we do not exclude rare or context-specific selection events, our results argue against pervasive somatic selection and highlight the importance of rigorous stratification when interpreting mtDNA mutational patterns. ### Competing Interest Statement The authors have declared no competing interest.
Aging, characterized by a series of functional declines correlated with advancing chronological age, has a significant mitochondrial DNA (mtDNA) component, with somatic mtDNA deletions playing a central role. In post-mitotic or slow-dividing cells like neurons and skeletal muscles, selfish mtDNA deletions clonally expand within a cell, ultimately leading to the deterioration and death of host cells and appearence of age-related phenotypes. Thus reducing the burden of somatic deletions could have far-reaching systemic benefits for the entire human body. Given the crucial role of direct nucleotide repeats in the formation of mitochondrial deletions, we hypothesize that minimizing these repeats in the human mitochondrial genome could enhance healthspan by decreasing somatic deletions. To investigate this hypothesis, we focus on the "common repeat", a 13-base pair perfect direct repeat sequence (ACCTCCCTCACCA) located at positions 8470-8482 and 13447-13459, respectively. This perfect repeat: (i) is highly prevalent, with its potential deleterious consequences affecting the majority of humans; (ii) represents one of the most fragile sites, highly prone to forming deletions; (iii) when disrupted, is associated with a decreased somatic deletion load and enhanced human healthspan; (iv) is likely to experience positive selection in the present or near future due to indirect fitness effects, such as the "grandmother effect", and direct fitness effects, such as (v) a decreased mutation rate. These observations support the argument that reducing the mtDNA somatic deletion load through targeted disruption of these repeats, or by using naturally occurring polymorphisms with disrupted repeats in mitochondrial medicine, could be an effective approach to increasing human longevity. ### Competing Interest Statement The authors have declared no competing interest.
Somatic mitochondrial DNA (mtDNA) mutations are prevalent in tumors, yet defining their biological significance remains challenging due to the intricate interplay between selective pressure, heteroplasmy, and cell state. Utilizing bulk whole-genome sequencing data from matched tumor and normal samples from two cohorts of pediatric cancer patients, we uncover differences in the accumulation of synonymous and nonsynonymous mtDNA mutations in pediatric leukemias, indicating distinct selective pressures. By integrating single-cell sequencing (SCS) with mathematical modeling and network-based systems biology approaches, we identify a correlation between the extent of cell-state changes associated with tumor-enriched mtDNA mutations and the selective pressures shaping their distribution among individual leukemic cells. Our findings also reveal an association between specific heteroplasmic mtDNA mutations and cellular responses that may contribute to functional heterogeneity among leukemic cells and influence their fitness. This study highlights the potential of SCS strategies for distinguishing between pathogenic and passenger somatic mtDNA mutations in cancer.
Serrano et al. (Serrano et al., 2024) use a high-fidelity somatic mtDNA mutation analysis in conplastic mice in which mtDNA was replaced with exogenous mtDNA of different mouse strains. Serrano reported apparent abundant somatic reversion mutations in the exogenous mtDNA that seemed to restore the original mito-nuclear match. If real, such a phenomenon would have important implications for health and genetics. In todays highly mixed human population, the pairing of potentially mismatched nuclear and mitochondrial genomes is widespread, so the proposed reversion mutagenesis should be commonplace. We demonstrate, however, that these reversion mutations are not real but originate from cross-contamination between samples and from NUMTs, the mtDNA pseudogenes located in the nuclear genome. ### Competing Interest Statement The authors have declared no competing interest.
The resilience of the mitochondrial genome (mtDNA) to a high mutational pressure depends, in part, on negative purifying selection in the germline. A paradigm in the field has been that such selection, at least in part, takes place in primordial germ cells (PGCs). Specifically, Floros et al. (Nature Cell Biology 20: 144-51) reported an increase in the synonymity of mtDNA mutations (a sign of purifying selection) between early-stage and late-stage PGCs. We re-analyzed Floros' et al. data and determined that their mutational dataset was significantly contaminated with single nucleotide variants (SNVs) derived from a nuclear sequence of mtDNA origin (NUMT) located on chromosome 5. Contamination was caused by co-amplification of the NUMT sequence by crossspecific PCR primers. Importantly, when we removed NUMT-derived SNVs, the evidence of purifying selection was abolished. In addition to bulk PGCs, Floros et al. reported the analysis of single-cell late-stage PGCs, which were amplified with different sets of PCR primers that cannot amplify the NUMT sequence. Accordingly, there were no NUMT-derived SNVs among single PGC mutations. Interestingly, single PGC mutations show a decrease of synonymity with increased intracellular mutant fraction. More specifically, nonsynonymous mutations show faster intracellular genetic drift towards higher mutant fraction than synonymous ones. This pattern is incompatible with predominantly negative selection. This suggests that germline selection of mtDNA mutations is a complex phenomenon and that the part of this process that takes place in PGCs may be predominantly positive. However counterintuitive, positive germline selection of detrimental mtDNA mutations has been reported previously and potentially may be evolutionarily advantageous.
To elucidate the primary factors shaping mitochondrial DNA (mtDNA) mutagenesis, we derived a comprehensive 192-component mtDNA mutational spectrum using 86,149 polymorphic synonymous mutations reconstructed from the CytB gene of 967 chordate species. The mtDNA spectrum analysis provided numerous findings on repair and mutation processes, breaking it down into three main signatures: (i) symmetrical, evenly distributed across both strands, mutations, induced by gamma DNA polymerase (about 50% of all mutations); (ii) asymmetrical, heavy-strand-specific, C>T mutations (about 30%); and (iii) asymmetrical, heavy-strand-specific A>G mutations, influenced by metabolic and age-specific factors (about 20%). We propose that both asymmetrical signatures are driven by single-strand specific damage coupled with inefficient base excision repair on the lagging (heavy) strand of mtDNA. Understanding the detailed mechanisms of this damage is crucial for developing strategies to reduce somatic mtDNA mutational load, which is vital for combating age-related diseases.### Competing Interest StatementThe authors have declared no competing interest.
The shift of the level of disease-causing mtDNA mutations (heteroplasmy) from mother to child is typically negatively correlated with the mother ′ s heteroplasmy (Hm). In other words, mothers with low Hm tend to have children with a higher mutation level (Hch) than their own. In contrast, mothers with high Hm typically see a decrease in heteroplasmy in their children. This peculiar trend has been commonly interpreted as a result of a descending germline selection profile, i.e., positive selection at low Hm, gradually turning negative at high Hm. Here we demonstrate, however, that the negative correlation is mostly driven by RTM, or ′ Regression To the Mean ′, a classical statistical bias. We further show that RTM can be nullified by using the average between the mother ′ s and child ′ s heteroplasmy, as a new variable, instead of the commonly used mother ′ s heteroplasmy in blood. Additionally, we demonstrate that mother/child average is a better proxy of the actual germline heteroplasmy. Moreover, the elimination of RTM revealed a previously hidden wave-shaped HS-profile (positive mother-to-child shift at intermediate average mother/child heteroplasmy, decreasing towards high and low average heteroplasmy). In confirmation of this finding, we show that simulations that involve both wave-shaped HS-profile and RTM, reproduce the observed patterns of inheritance of mtDNA mutations in unprecedented detail. From the health care perspective, the uncovering of the wave-shaped HS-profile (and the removal of the RTM bias) are crucial for millions of families affected by mtDNA disease. From the fundamental perspective, the wave-shaped profile offers a novel understanding of the germline dynamics of mtDNA and a novel potential mechanism that prevents the spread of detrimental mtDNA mutations in the population.
Levenshtein distance is a commonly used edit distance metric, typically applied in language processing, and to a lesser extent, in molecular biology analysis. Biological nucleic acid sequences are often embedded in longer sequences and are subject to insertion and deletion errors that introduce frameshift during sequencing. These frameshift errors are due to string context and should not be counted as true biological errors. Sequence-Levenshtein distance is a modification to Levenshtein distance that is permissive of frameshift error without additional penalty. However, in a biological context Levenshtein distance needs to accommodate both frameshift and weighted errors, which Sequence-Levenshtein distance cannot do. Errors are weighted when they are associated with a numerical cost that corresponds to their frequency of appearance. Here, we describe a modification that allows the use of Levenshtein distance and Sequence-Levenshtein distance to appropriately accommodate penalty-free frameshift between embedded sequences and correctly weight specific error types.
Background Aging in postmitotic tissues is associated with clonal expansion of somatic mitochondrial deletions, the origin of which is not well understood. Such deletions are often flanked by direct nucleotide repeats, but this alone does not fully explain their distribution. Here, we hypothesized that the close proximity of direct repeats on single-stranded mitochondrial DNA (mtDNA) might play a role in the formation of deletions. Results By analyzing human mtDNA deletions in the major arc of mtDNA, which is single-stranded during replication and is characterized by a high number of deletions, we found a non-uniform distribution with a “hot spot” where one deletion breakpoint occurred within the region of 6–9 kb and another within 13–16 kb of the mtDNA. This distribution was not explained by the presence of direct repeats, suggesting that other factors, such as the spatial proximity of these two regions, can be the cause. In silico analyses revealed that the single-stranded major arc may be organized as a large-scale hairpin-like loop with a center close to 11 kb and contacting regions between 6–9 kb and 13–16 kb, which would explain the high deletion activity in this contact zone. The direct repeats located within the contact zone, such as the well-known common repeat with a first arm at 8470–8482 bp (base pair) and a second arm at 13,447–13,459 bp, are three times more likely to cause deletions compared to direct repeats located outside of the contact zone. A comparison of age- and disease-associated deletions demonstrated that the contact zone plays a crucial role in explaining the age-associated deletions, emphasizing its importance in the rate of healthy aging. Conclusions Overall, we provide topological insights into the mechanism of age-associated deletion formation in human mtDNA, which could be used to predict somatic deletion burden and maximum lifespan in different human haplogroups and mammalian species.
Every cell in our body contains a vibrant population of mitochondria, or, more precisely, of mitochondrial DNA molecules (mtDNAs). Just like members of any population mtDNAs multiply (by replication) and ‘die’ (i.e., are removed, either by degradation or by distribution into the sister cell in mitosis). An intriguing question is whether all mitochondria in this population are equal, especially whether some are responsible primarily for reproduction and some - for empowering the various jobs of the mitochondrion, oxidative phosphorylation in the first place. Importantly, because mtDNA is highly damaged such a separation of responsibilities could help greatly reduce the conversion of DNA damage into real inheritable mutations. An unexpected twist in the resolution of this problem has been brought about by a recent high-precision analysis of mtDNA mutations (Sanchez-Contreras et al. 2023). They discovered that certain transversion mutations, unlike more common transitions, are not accumulating with age in mice. We argue that this observation requires the existence of a permanent replicating subpopulation/lineage of mtDNA molecules, which are protected from DNA damage, a.k.a. the ‘stem’ mtDNA. This also implies the existence of its antipode i.e., the ‘worker’ mtDNA, which empowers OSPHOS, sustains damage and rarely replicates. The analysis of long HiFi reads of mtDNA performed by PacBio closed circular sequencing confirms this assertion.
A large-scale study of mutations in mitochondrial DNA has revealed a subset that do not accumulate with age.
The mutational spectrum of the mitochondrial DNA (mtDNA) does not resemble any of the known mutational signatures of the nuclear genome and variation in mtDNA mutational spectra between different organisms is still incomprehensible. Since mitochondria are responsible for aerobic respiration, it is expected that mtDNA mutational spectrum is affected by oxidative damage. Assuming that oxidative damage increases with age, we analyse mtDNA mutagenesis of different species in regards to their generation length. Analysing, (i) dozens of thousands of somatic mtDNA mutations in samples of different ages (ii) 70053 polymorphic synonymous mtDNA substitutions reconstructed in 424 mammalian species with different generation lengths and (iii) synonymous nucleotide content of 650 complete mitochondrial genomes of mammalian species we observed that the frequency of A(H) > G(H) substitutions (H: heavy strand notation) is twice bigger in species with high versus low generation length making their mtDNA more A(H) poor and G(H) rich. Considering that A(H) > G(H) substitutions are also sensitive to the time spent single-stranded (TSSS) during asynchronous mtDNA replication we demonstrated that A(H) > G(H) substitution rate is a function of both species-specific generation length and position-specific TSSS. We propose that A(H) > G(H) is a mitochondria-specific signature of oxidative damage associated with both aging and TSSS.
The A-to-G point mutation at position 3243 in the human mitochondrial genome (m.3243A>G) is the most common pathogenic mtDNA variant responsible for disease in humans. It is widely accepted that m.3243A>G levels decrease in blood with age, and an age correction representing ∼2% annual decline is often applied to account for this change in mutation level. Here we report that recent data indicate the dynamics of m.3243A>G are more complex and depend on the mutation level in blood in a bi-phasic way. Consequently, the traditional 2% correction, which is adequate ‘on average’, creates opposite predictive biases at high and low mutation levels. Unbiased age correction is needed to circumvent these drawbacks of the standard model. We propose to eliminate both biases by using an approach where age correction depends on mutation level in a biphasic way to account for the dynamics of m.3243A>G in blood. The utility of this approach was further tested in estimating germline selection of m.3243A>G. The biphasic approach permitted us to uncover patterns consistent with the possibility of positive selection for m.3243A>G. Germline selection of m.3243A>G shows an ‘arching’ profile by which selection is positive at intermediate mutant fractions and declines at high and low mutant fractions. We conclude that use of this biphasic approach will greatly improve the accuracy of modelling changes in mtDNA mutation frequencies in the germline and in somatic cells during aging.
Background Third-generation sequencing offers some advantages over next-generation sequencing predecessors, but with the caveat of harboring a much higher error rate. Clustering-related sequences is an essential task in modern biology. To accurately cluster sequences rich in errors, error type and frequency need to be accounted for. Levenshtein distance is a well-established mathematical algorithm for measuring the edit distance between words and can specifically weight insertions, deletions and substitutions. However, there are drawbacks to using Levenshtein distance in a biological context and hence has rarely been used for this purpose. We present novel modifications to the Levenshtein distance algorithm to optimize it for clustering error-rich biological sequencing data. Results We successfully introduced a bidirectional frameshift allowance with end-user determined accommodation caps combined with weighted error discrimination. Furthermore, our modifications dramatically improved the computational speed of Levenstein distance. For simulated ONT MinION and PacBio Sequel datasets, the average clustering sensitivity for 3GOLD was 41.45% (S.D. 10.39) higher than Sequence-Levenstein distance, 52.14% (S.D. 9.43) higher than Levenshtein distance, 55.93% (S.D. 8.67) higher than Starcode, 42.68% (S.D. 8.09) higher than CD-HIT-EST and 61.49% (S.D. 7.81) higher than DNACLUST. For biological ONT MinION data, 3GOLD clustering sensitivity was 27.99% higher than Sequence-Levenstein distance, 52.76% higher than Levenshtein distance, 56.39% higher than Starcode, 48% higher than CD-HIT-EST and 70.4% higher than DNACLUST. Conclusion Our modifications to Levenshtein distance have improved its speed and accuracy compared to the classic Levenshtein distance, Sequence-Levenshtein distance and other commonly used clustering approaches on simulated and biological third-generation sequenced datasets. Our clustering approach is appropriate for datasets of unknown cluster centroids, such as those generated with unique molecular identifiers as well as known centroids such as barcoded datasets. A strength of our approach is high accuracy in resolving small clusters and mitigating the number of singletons.