Phylogenetic tree construction from homologous sequences is a fundamental approach for studying evolutionary relationships. However, in cases where sequence similarity is too low to generate reliable alignments, comparison of protein structure provides an alternative means for inferring deep evolutionary relationships. MD-phylogeny is a method that integrates structural comparison with molecular dynamics (MD) simulations to generate robust phylogenies for protein datasets within the twilight zone of sequence similarity. This approach uses well-established structural superposition methods for building trees from structure, with a method for generating structural variants using MD simulations, from which confidence scores analogous to bootstrapping can be derived.This protocol provides a step-by-step guide to performing MD-phylogeny, covering dataset selection, structural comparisons, tree construction, and the generation of statistical support through MD simulations.
Large genomes such as the human genome are pervasively transcribed yet encode relatively few unambiguously functional elements. This has led to debate over whether pervasive transcription is indicative of large suites of uncharacterized functional elements or is simply background noise. Here, we used a deep-learning model to estimate background transcription in the human genome as a way of distinguishing between these two hypotheses. We applied the model to randomized (reversed or shuffled) versions of the human genome and found that transcription is predicted to be sparse across all randomization methods, initiating with at least four-fold lower frequencies than in the native human genome. This relatively low level of background transcription from the human genome suggests that most transcription is not a consequence of background noise, thus it requires other explanations. We find that randomizing only interspersed repeats in human genome has little impact on predicted transcription, suggesting that transcription of mobile elements does not explain the excess transcription in the human genome. Instead, most transcriptional events may derive from functional noncoding RNA transcripts, some general requirement for extensive transcription initiation/elongation, and/or mutational biases leading to the frequent appearance of transcription initiation sites by chance.
In biology, changes to a DNA sequence can impact protein sequence but changes to protein sequences (phenotype) do not flow back into DNA (genotype). A system with bidirectional information flow (i.e., both translation and 'reverse translation') remains a theoretical possibility for an independent origin of life or an artificial biosystem, but the recent development of digital data storage in DNA does just this: changes made to a digital file can be written back into DNA, meaning changes to 'phenotype' can be written back to 'genotype'. To explore the evolutionary properties of such a system, we created an artificial system where synthetic DNA serves as genotype and music as phenotype. Audio can be output from a DNA sequence, then recorded and written to DNA as 'codons', enabling bidirectional information flow (DNA→music and music→DNA). Our results show that the mutation rate in a bidirectional system is much higher than for unidirectional information flow, and that, under reverse translation there is no mechanism for preservation of codon choice across generations. This has the effect of eliminating the impact of spontaneous synonymous mutations, a key benefit of a redundant genetic code. As a result, non-synonymous mutations are the only DNA-level changes that are transmitted across generations, and, as non-synonymous mutations can emerge at both 'genotypic' and 'phenotypic' levels, these occur at a two-fold higher frequency than in a unidirectional system. Our system holds some practical insight. First, for DNA read/write systems, it may be wise to avoid designing systems with 'de novo reverse translation' because the opportunities for mutation are higher; tracking genotype information from the preceding generation to guide this process may reduce error. Second, our system helps clarify how a 'Lamarckian' biological system might operate. We conclude that, were a 'Lamarckian' system of inheritance a feature of early genetic systems, it would likely have been short lived as the high frequency of mutation would risk driving the system to extinction. A system based on unidirectional information flow thus appears superior as there are fewer opportunities for mutational error.
Despite fundamental advances in cellular, molecular and genome biology, there is still surprisingly little consensus concerning the evolutionary origins of the eukaryote cell. While it is clear that the mitochondrion (responsible for generating much of the energy requirements of the eukaryote cell) has evolved from an endosymbiont cell of bacterial origin, the recent literature has borne witness to a tidal wave of speculative theories regarding the nature of the cell in which this bacterium took up residence. David Penny and I recently argued that much of this confusion can be avoided if models are grounded in known biological processes, and if speculation is tempered by formulating testable hypotheses. The most fanciful hypotheses are an inevitable casualty of a pragmatic approach, but what remains is a productive framework wherein biologically plausible alternatives can be evaluated without the need to invoke ad hoc events or processes, such as biological ‘big bangs’ or hitherto unobserved cell biological phenomena.
Duplication is a major route for the emergence of new gene functions. However, the emergence of new gene functions via this route may be reduced in prokaryotes, as redundant genes are often rapidly purged. In lineages with compact, streamlined genomes, it thus appears challenging for novel function to emerge via duplication and divergence. A further pressure contributing to gene loss occurs under Black Queen dynamics, as cheaters that lose the capacity to produce a public good can instead acquire it from neighbouring producers. We propose that Black Queen dynamics can favour the emergence of new function because, under an emerging Black Queen dynamic, there is high gene redundancy spread across a community of interacting cells. Using computational modelling, we demonstrate that new gene functions can emerge under Black Queen dynamics. This result holds even if there is deletion bias due to low duplication rates and selection against redundant gene copies resulting from the high cost associated with carrying a locus. However, when the public good production costs are high, Black Queen dynamics impede the fixation of new functions. Our results expand the mechanisms by which new gene functions can emerge in prokaryotic systems.
Theoretical biologist who ‘tamed’ mathematicians and tested the theory of evolution.
Protein structure is more conserved than protein sequence, and therefore may be useful for phylogenetic inference beyond the “twilight zone” where sequence similarity is highly decayed. Until recently, structural phylogenetics was constrained by the lack of solved structures for most proteins, and the reliance on phylogenetic distance methods which made it difficult to treat inference and uncertainty statistically. AlphaFold has mostly overcome the first problem by making structural predictions readily available. We address the second problem by redeploying a structural alphabet recently developed for Foldseek, a highly-efficient deep homology search program. For each residue in a structure, Foldseek identifies a tertiary interaction closest-neighbor residue in the structure, and classifies it into one of twenty “3Di” states. We test the hypothesis that 3Dis can be used as standard phylogenetic characters using a dataset of 53 structures from the ferritin-like superfamily. We performed 60 IQtree Maximum Likelihood runs to compare structure-free, PDB, and AlphaFold analyses, and default versus custom model sets that include a 3DI-specific rate matrix. Analyses that combine amino acids, 3Di characters, partitioning, and custom models produce the closest match to the structural distances tree of [Malik et al. (2020)][1], avoiding the long-branch attraction errors of structure-free analyses. Analyses include standard ultrafast bootstrapping confidence measures, and take minutes instead of weeks to run on desktop computers. These results suggest that structural phylogenetics could soon be routine practice in protein phylogenetics, allowing the re-exploration of many fundamental phylogenetic problems.### Competing Interest StatementThe authors have declared no competing interest. [1]: #ref-43
Life requires ribonucleotide reduction for de novo synthesis of deoxyribonucleotides. As ribonucleotide reduction has on occasion been lost in parasites and endosymbionts, which are instead dependent on their host for deoxyribonucleotide synthesis, it should in principle be possible to knock this process out if growth media are supplemented with deoxyribonucleosides. We report the creation of a strain of Escherichia coli where all three ribonucleotide reductase operons have been deleted following introduction of a broad spectrum deoxyribonucleoside kinase from Mycoplasma mycoides. Our strain shows slowed but substantial growth in the presence of deoxyribonucleosides. Under limiting deoxyribonucleoside levels, we observe a distinctive filamentous cell morphology, where cells grow but do not appear to divide regularly. Finally, we examined whether our lines can adapt to limited supplies of deoxyribonucleosides, as might occur in the switch from de novo synthesis to dependence on host production during the evolution of parasitism or endosymbiosis. Over the course of an evolution experiment, we observe a 25-fold reduction in the minimum concentration of exogenous deoxyribonucleosides necessary for growth. Genome analysis reveals that several replicate lines carry mutations in deoB and cdd. deoB codes for phosphopentomutase, a key part of the deoxyriboaldolase pathway, which has been hypothesised as an alternative to ribonucleotide reduction for deoxyribonucleotide synthesis. Rather than complementing the loss of ribonucleotide reduction, our experiments reveal that mutations appear that reduce or eliminate the capacity for this pathway to catabolise deoxyribonucleotides, thus preventing their loss via central metabolism. Mutational inactivation of both deoB and cdd is also observed in a number of obligate intracellular bacteria that have lost ribonucleotide reduction. We conclude that our experiments recapitulate key evolutionary steps in the adaptation to life without ribonucleotide reduction.
Abstract Summary Protein structures carry signal of common ancestry and can therefore aid in reconstructing their evolutionary histories. To expedite the structure-informed inference process, a web server, Structome, has been developed that allows users to rapidly identify protein structures similar to a query protein and to assemble datasets useful for structure-based phylogenetics. Structome was created by clustering ∼94% of the structures in RCSB PDB using 90% sequence identity and representing each cluster by a centroid structure. Structure similarity between centroid proteins was calculated, and annotations from PDB, SCOP, and CATH were integrated. To illustrate utility, an H3 histone was used as a query, and results show that the protein structures returned by Structome span both sequence and structural diversity of the histone fold. Additionally, the pre-computed nexus-formatted distance matrix, provided by Structome, enables analysis of evolutionary relationships between proteins not identifiable using searches based on sequence similarity alone. Our results demonstrate that, beginning with a single structure, Structome can be used to rapidly generate a dataset of structural neighbours and allows deep evolutionary history of proteins to be studied. Availability and Implementation Structome is available at: https://structome.bii.a-star.edu.sg.
The nuclear pore is structurally conserved across eukaryotes as are many of the pore's constituent proteins. The transmembrane nuclear pore proteins GP210 and NDC1 span the nuclear envelope holding the nuclear pore in place. Orthologues of GP210 and NDC1 in Arabidopsis were investigated through characterisation of T-DNA insertional mutants. While the T-DNA insert into GP210 reduced expression of the gene, the insert in the NDC1 gene resulted in increased expression in both the ndc1 mutant as well as the ndc1/gp210 double mutant. The ndc1 and gp210 individual mutants showed little phenotypic difference from wild-type plants, but the ndc1/gp210 mutant showed a range of phenotypic effects. As with many plant nuclear pore protein mutants, these effects included non-nuclear phenotypes such as reduced pollen viability, reduced growth and glabrous leaves in mature plants. Importantly, however, ndc1/gp210 exhibited nuclear-specific effects including modifications to nuclear shape in different cell types. We also observed functional changes to nuclear transport in ndc1/gp210 plants, with low levels of cytoplasmic fluorescence observed in cells expressing nuclear-targeted GFP. The lack of phenotypes in individual insertional lines, and the relatively mild phenotype suggests that additional transmembrane nucleoporins, such as the recently-discovered CPR5, likely compensate for their loss.
Protein structures carry signal of common ancestry and can therefore aid in reconstructing their evolutionary histories. To expedite the structure-informed inference process, a web server, Structome, has been developed, that allows users to rapidly identify protein structures similar to a query protein and to assemble datasets useful for structure-based phylogenetics. Structome was created by clustering ∼ 94% of the structures in RCSB PDB using 90% sequence identity and representing each cluster by a centroid structure. Structure similarity between centroid proteins was calculated, and annotations from PDB, SCOP and CATH were integrated. To illustrate utility, an H3 histone was used as a query, and results show that the protein structures returned by Structome span both sequence and structural diversity of the histone fold. Additionally, the pre-computed nexus-formated distance matrix, provided by Structome, enables analysis of evolutionary relationships between proteins not identifiable using searches based on sequence similarity alone. Our results demonstrate that, beginning with a single structure, Structome can be used to rapidly generate a dataset of structural neighbours and allows deep evolutionary history of proteins to be studied. Structome is available at: https://structome.bii.a-star.edu.sg
All life requires ribonucleotide reduction for de novo synthesis of deoxyribonucleotides. A handful of obligate intracellular species are known to lack ribonucleotide reduction and are instead dependent on their host for deoxyribonucleotide synthesis. As ribonucleotide reduction has on occasion been lost in obligate intracellular parasites and endosymbionts, we reasoned that it should in principle be possible to knock this process out entirely under conditions where deoxyribonucleotides are present in the growth media. We report here the creation of a strain of E. coli where all three ribonucleotide reductase operons have been fully deleted. Our strain is able to grow in the presence of deoxyribonucleosides and shows slowed but substantial growth. Under limiting deoxyribonucleoside levels, we observe a distinctive filamentous cell morphology, where cells grow but do not appear to divide regularly. Finally, we examined whether our lines are able to adapt to limited supplies of deoxyribonucleosides, as might occur in the evolutionary switch from de novo synthesis to dependence on host production during the evolution of parasitism or endosymbiosis. Over the course of an evolution experiment, we observe a 25-fold reduction in the minimum concentration of exogenous deoxyribonucleosides necessary for growth. Genome analysis of replicate lines reveals that several lines carry mutations in deoB and cdd. deoB codes for phosphopentomutase, a key part of the deoxyriboaldolase pathway, which has been hypothesised as an alternative to ribonucleotide reduction for deoxyribonucleotide synthesis. Rather than synthesis via this pathway complementing the loss of ribonucleotide reduction, our experiments reveal that mutations appear that reduce or eliminate the capacity for this pathway to catabolise deoxyribonucleotides, thus preventing their loss via central metabolism. Mutational inactivation of both deoB and cdd is also observed in a number of obligate intracellular bacteria that have lost ribonucleotide reduction. We conclude that our experiments recapitulate key evolutionary steps in the adaptation to intracellular life without ribonucleotide reduction.
Masting, the synchronous, highly variable flowering across years by a population of perennial plants, has been reported to be precipitated by various factors including nitrogen levels, drought conditions, and spring and summer temperatures. However, the molecular mechanism leading to the initiation of flowering in masting plants in particular years remains largely unknown, despite the potential impact of climate change on masting phenology. We studied genes controlling flowering in the alpine snow tussock Chionochloa pallens (Poaceae), a strongly masting perennial grass. We used a range of in situ and manipulated plants to obtain leaf samples from tillers (shoots) which subsequently remained vegetative or flowered. Here, we show that a novel orthologue of TERMINAL FLOWER 1 (TFL1; normally a repressor of flowering in other species) promotes the induction of flowering in C. pallens (hence Anti-TFL1), a conclusion supported by structural, functional and expression analyses. Global transcriptomic analysis indicated differential expression of CpTPS1, CpGA20ox1, CpREF6 and CpHDA6, emphasizing the role of endogenous cues and epigenetic regulation in terms of responsiveness of plants to initiate flowering. Our molecular-based study provides insights into the cellular mechanism of flowering in masting plants and will supplement ecological and statistical models to predict how masting will respond to global climate change.
Mast flowering (or masting) is synchronous, highly variable flowering among years in populations of perennial plants. Despite having widespread consequences for seed consumers, endangered fauna and human health, masting is hard to predict. While observational studies show links to various weather patterns in different plant species, the mechanism(s) underpinning the regulation of masting is still not fully explained. We studied floral induction in Celmisia lyallii (Asteraceae), a mast flowering herbaceous alpine perennial, comparing gene expression in flowering and nonflowering plants. We performed translocation experiments to induce the floral transition in C. lyallii plants followed by both global and targeted expression analysis of flowering-pathway genes. Differential expression analysis showed elevated expression of ClSOC1 and ClmiR172 (promoters of flowering) in leaves of plants that subsequently flowered, in contrast to elevated expression of ClAFT and ClTOE1 (repressors of flowering) in leaves of plants that did not flower. The warm summer conditions that promoted flowering led to differential regulation of age and hormonal pathway genes, including ClmiR172 and ClGA20ox2, known to repress the expression of floral repressors and permit flowering. Upregulated expression of epigenetic modifiers of floral promoters also suggests that plants may maintain a novel "summer memory" across years to induce flowering. These results provide a basic mechanistic understanding of floral induction in masting plants and evidence of their ability to imprint various environmental cues to synchronize flowering, allowing us to better predict masting events under climate change.
Masting, the synchronous highly variable flowering across years by a population of perennial plants, has been shown to be precipitated by many factors including nitrogen levels, drought conditions, spring and summer temperatures. However, the molecular mechanism leading to the initiation of flowering in masting plants in particular years remains largely unknown, despite the potential impact of climate change on masting phenology. We studied genes controlling flowering in Chionochloa pallens, a strongly masting perennial grass. We used a range of in situ and manipulated plants to obtain leaf samples from tillers (shoots) which subsequently remained vegetative or flowered. Here, we show that a novel orthologue of TERMINAL FLOWER 1 (TFL1; normally a repressor of flowering in other species) promotes the induction of flowering in C. pallens (hence Anti-TFL1), a conclusion supported by structural, functional and expression analyses. Global transcriptomic analysis indicated differential expression of CpTPS1, CpGA20ox1, CpREF6 and CpHDA6, emphasising the role of endogenous cues and epigenetic regulation in terms of responsiveness of plants to initiate flowering. Our molecular-based study has provided insights into the cellular mechanism of flowering in masting plants and will supplement ecological and statistical models to predict how masting will respond to global climate change.
There is general agreement that bacteria, archaea, and eukarya share common ancestry. However, tracing back extant lineages to reconstruct the ancestral gene set of the three domains has proven to be non-trivial, as there is little unambiguous signal this far back in time. In this chapter, I explain the basic principles behind reconstruction of the Last Universal Common Ancestor (LUCA) and summarise a few of the challenges associated with reconstruction. Finally, I consider whether a mid-resolution LUCA might be the most achievable goal, particularly from the perspective of the classes of chemistry available to early life.
EDITORIAL article Front. Microbiol., 30 August 2021 | https://doi.org/10.3389/fmicb.2021.751416
Background Trichomonas vaginalis , the causative agent of a prevalent urogenital infection in humans, is an evolutionarily divergent protozoan. Protein-coding genes in T. vaginalis are largely controlled by two core promoter elements, producing mRNAs with short 5′ UTRs. The specific mechanisms adopted by T. vaginalis to fine-tune the translation efficiency (TE) of mRNAs remain largely unknown. Results Using both computational and experimental approaches, this study investigated two key factors influencing TE in T. vaginalis : codon usage and mRNA secondary structure. Statistical dependence between TE and codon adaptation index (CAI) highlighted the impact of codon usage on mRNA translation in T. vaginalis . A genome-wide interrogation revealed that low structural complexity at the 5′ end of mRNA followed closely by a highly structured downstream region correlates with TE variation in this organism. To validate these findings, a synthetic library of 15 synonymous iLOV genes was created, representing five mRNA folding profiles and three codon usage profiles. Fluorescence signals produced by the expression of these synonymous iLOV genes in T. vaginalis were consistent with and validated our in silico predictions. Conclusions This study demonstrates the role of codon usage bias and mRNA secondary structure in TE of T. vaginalis mRNAs, contributing to a better understanding of the factors that influence, and possibly regulate, gene expression in this human pathogen.
For evaluating the deepest evolutionary relationships among proteins, sequence similarity is too low for application of sequence-based homology search or phylogenetic methods. In such cases, comparison of protein structures, which are often better conserved than sequences, may provide an alternative means of uncovering deep evolutionary signal. Although major protein structure databases such as SCOP and CATH hierarchically group protein structures, they do not describe the specific evolutionary relationships within a hierarchical level. Structural phylogenies have the potential to fill this gap. However, it is difficult to assess evolutionary relationships derived from structural phylogenies without some means of assessing confidence in such trees. We therefore address two shortcomings in the application of structural data to deep phylogeny. First, we examine whether phylogenies derived from pairwise structural comparisons are sensitive to differences in protein length and shape. We find that structural phylogenetics is best employed where structures have very similar lengths, and that shape fluctuations generated during molecular dynamics simulations impact pairwise comparisons, but not so drastically as to eliminate evolutionary signal. Second, we address the absence of statistical support for structural phylogeny. We present a method for assessing confidence in a structural phylogeny using shape fluctuations generated via molecular dynamics or Monte Carlo simulations of proteins. Our approach will aid the evolutionary reconstruction of relationships across structurally defined protein superfamilies. With the Protein Data Bank now containing in excess of 158,000 entries (December 2019), we predict that structural phylogenetics will become a useful tool for ordering the protein universe.