Antiviral defense in ecdysozoan invertebrates requires Dicer with a helicase domain capable of ATP hydrolysis. But despite well-conserved ATPase motifs, human Dicer is incapable of ATP hydrolysis, consistent with a muted role in antiviral defense. To investigate this enigma, we used ancestral protein reconstruction to resurrect Dicer’s helicase in animals and trace the evolutionary trajectory of ATP hydrolysis. Biochemical assays indicated ancient Dicer possessed ATPase function, that like extant invertebrate Dicers, is stimulated by dsRNA. Analyses revealed that dsRNA stimulates ATPase activity by increasing ATP affinity, reflected in Michaelis constants. Deuterostome Dicer-1 ancestor, while exhibiting lower dsRNA affinity, retained some ATPase activity; importantly, ATPase activity was undetectable in the vertebrate Dicer-1 ancestor, which had even lower dsRNA affinity. Reverting residues in the ATP hydrolysis pocket was insufficient to rescue hydrolysis, but additional substitutions distant from the pocket rescued vertebrate Dicer-1’s ATPase function. Our work suggests Dicer lost ATPase function in the vertebrate ancestor due to loss of ATP affinity, involving motifs distant from the active site, important for coupling dsRNA binding to the active conformation. By competing with Dicer for viral dsRNA, RIG-I-like receptors important for interferon signaling may have allowed or actively caused loss of ATPase function.
The evolution of biological nitrogen fixation, uniquely catalyzed by nitrogenase enzymes, has been one of the most consequential biogeochemical innovations over life's history. Though understanding the early evolution of nitrogen fixation has been a longstanding goal from molecular, biogeochemical, and planetary perspectives, its origins remain enigmatic. In this study, we reconstructed the evolutionary histories of nitrogenases, as well as homologous maturase proteins that participate in the assembly of the nitrogenase active-site cofactor but are not able to fix nitrogen. We combined phylogenetic and ancestral sequence inference with an analysis of predicted functionally divergent sites between nitrogenases and maturases to infer the nitrogen-fixing capabilities of their shared ancestors. Our results provide phylogenetic constraints to the emergence of nitrogen fixation and are consistent with a model wherein nitrogenases emerged from maturase-like predecessors. Though the precise functional role of such a predecessor protein remains speculative, our results highlight evolutionary contingency as a significant factor shaping the evolution of a biogeochemically essential enzyme.
Equine recurrent uveitis (ERU) is a painful and debilitating autoimmune disease and represents the only spontaneous model of human recurrent uveitis (RU). Despite the efficacy of existing treatments, RU remains a leading cause of visual handicap in horses and humans. Cytokines, which utilize Janus kinase 2 (Jak2) for signaling, drive the inflammatory processes in ERU that promote blindness. Notably, suppressor of cytokine signaling 1 (SOCS1), which naturally limits the activation of Jak2 through binding interactions, is often deficient in autoimmune disease patients. Significantly, we previously showed that topical administration of a SOCS1 peptide mimic (SOCS1-KIR) mitigated induced rodent uveitis. In this pilot study, we test the potential to translate the therapeutic efficacy observed in experimental rodent uveitis to equine patient disease. Through bioinformatics and peptide binding assays we demonstrate putative binding of the SOCS1-KIR peptide to equine Jak2. We also show that topical, or intravitreal injection of SOCS1-KIR was well tolerated within the equine eye through physical and ophthalmic examinations. Finally, we show that topical SOCS1-KIR administration was associated with significant clinical ERU improvement. Together, these results provide a scientific rationale, and supporting experimental evidence for the therapeutic use of a SOCS1 mimetic peptide in RU.
AbstractMotivationDetecting subtle biologically relevant patterns in protein sequences often requires the construction of a large and accurate multiple sequence alignment (MSA). Methods for constructing MSAs are usually evaluated using benchmark alignments, which, however, typically contain very few sequences and are therefore inappropriate when dealing with large numbers of proteins.ResultseCOMPASS addresses this problem using a statistical measure of relative alignment quality based on direct coupling analysis (DCA): to maintain protein structural integrity over evolutionary time, substitutions at one residue position typically result in compensating substitutions at other positions. eCOMPASS computes the statistical significance of the congruence between high scoring directly coupled pairs and 3D contacts in corresponding structures, which depends upon properly aligned homologous residues. We illustrate eCOMPASS using both simulated and real MSAs.Availability and implementationThe eCOMPASS executable, C++ open source code and input data sets are available at https://www.igs.umaryland.edu/labs/neuwald/software/compassSupplementary informationSupplementary data are available at Bioinformatics online.
ABSTRACTRNA interference (RNAi) plays important roles in organism development through post-transcriptional regulation of specific target mRNAs. Target specificity is largely controlled by base-pair complementarity between micro-RNA (miRNA) regulatory elements and short regions of the target mRNA. The pattern of miRNA production in a cell interacts with the cell’s mRNA transcriptome to generate a specific network of post-transcriptional regulation that can play critical roles in cellular metabolism, differentiation, tissue/organ development and developmental timing. In plants, miRNA production is orchestrated in the nucleus by a suite of proteins that control transcription of the pri-miRNA gene, post-transcriptional processing and nuclear export of the mature miRNA. In the model plant, Arabidopsis thaliana, post-transcriptional processing of miRNAs is controlled by a pair of physically-interacting proteins, HYL1 and DCL1. However, the evolutionary history of the HYL1-DCL1 interaction is unknown, as is its structural basis. Here we use ancestral sequence reconstruction and functional characterization of ancestral HYL1 in vitro and in vivo to better understand the origin and evolution of the HYL1-DCL1 interaction and its impact on miRNA production and plant development. We found the ancestral plant HYL1 evolved high affinity for both double-stranded RNA (dsRNA) and its DCL1 partner very early in plant evolutionary history, before the divergence of mosses from seed plants (~500 Ma), and these high-affinity interactions remained largely conserved throughout plant evolutionary history. Structural modeling and molecular binding experiments suggest that the second of two double-stranded RNA-binding motifs (DSRMs) in HYL1 may interact tightly with the first of two C-terminal DCL1 DSRMs to mediate the HYL1-DCL1 physical interaction necessary for efficient miRNA production. Transgenic expression of the nearly 200 Ma-old ancestral flowering-plant HYL1 in A. thaliana was sufficient to rescue many key aspects of plant development disrupted by HYL1− knockout and restored near-native miRNA production, suggesting that the functional partnership of HYL1-DCL1 originated very early in and was strongly conserved throughout the evolutionary history of terrestrial plants. Overall, our results are consistent with a model in which miRNA-based gene regulation evolved as part of a conserved plant ‘developmental toolkit’; its role in generating developmental novelty is probably related to the relatively rapid evolution of miRNA genes.
The nitrogenase metalloenzyme family, essential for supplying fixed nitrogen to the biosphere, is one of life's key biogeochemical innovations. The three forms of nitrogenase differ in their metal dependence, each binding either a FeMo-, FeV-, or FeFe-cofactor where the reduction of dinitrogen takes place. The history of nitrogenase metal dependence has been of particular interest due to the possible implication that ancient marine metal availabilities have significantly constrained nitrogenase evolution over geologic time. Here, we reconstructed the evolutionary history of nitrogenases, and combined phylogenetic reconstruction, ancestral sequence inference, and structural homology modeling to evaluate the potential metal dependence of ancient nitrogenases. We find that active-site sequence features can reliably distinguish extant Mo-nitrogenases from V- and Fe-nitrogenases and that inferred ancestral sequences at the deepest nodes of the phylogeny suggest these ancient proteins most resemble modern Mo-nitrogenases. Taxa representing early-branching nitrogenase lineages lack one or more biosynthetic nifE and nifN genes that both contribute to the assembly of the FeMo-cofactor in studied organisms, suggesting that early Mo-nitrogenases may have utilized an alternate and/or simplified pathway for cofactor biosynthesis. Our results underscore the profound impacts that protein-level innovations likely had on shaping global biogeochemical cycles throughout the Precambrian, in contrast to organism-level innovations that characterize the Phanerozoic Eon.
Ancestral sequence reconstruction (ASR) uses an alignment of extant protein sequences, a phylogeny describing the history of the protein family and a model of the molecular-evolutionary process to infer the sequences of ancient proteins, allowing researchers to directly investigate the impact of sequence evolution on protein structure and function. Like all statistical inferences, ASR can be sensitive to violations of its underlying assumptions. Previous studies have shown that, whereas phylogenetic uncertainty has only a very weak impact on ASR accuracy, uncertainty in the protein sequence alignment can more strongly affect inferred ancestral sequences. Here, we show that errors in sequence alignment can produce errors in ASR across a range of realistic and simplified evolutionary scenarios. Importantly, sequence reconstruction errors can lead to errors in estimates of structural and functional properties of ancestral proteins, potentially undermining the reliability of analyses relying on ASR. We introduce an alignment-integrated ASR approach that combines information from many different sequence alignments. We show that integrating alignment uncertainty improves ASR accuracy and the accuracy of downstream structural and functional inferences, often performing as well as highly accurate structure-guided alignment. Given the growing evidence that sequence alignment errors can impact the reliability of ASR studies, we recommend that future studies incorporate approaches to mitigate the impact of alignment uncertainty. Probabilistic modeling of insertion and deletion events has the potential to radically improve ASR accuracy when the model reflects the true underlying evolutionary history, but further studies are required to thoroughly evaluate the reliability of these approaches under realistic conditions.
The nitrogenase metalloenzyme family, essential for supplying fixed nitrogen to the biosphere, is one of life s key biogeochemical innovations. The three isozymes of nitrogenase differ in their metal dependence, each binding either a FeMo-, FeV-, or FeFe-cofactor for the reduction of nitrogen. The history of nitrogenase metal dependence has been of particular interest due to the possible implication that ancient marine metal availabilities have significantly constrained nitrogenase evolution over geologic time. Here, we combine phylogenetics and ancestral sequence reconstruction, a method by which inferred, historical protein sequence information can be linked to functional molecular properties, to reconstruct the metal dependence of ancient nitrogenases. Inferred ancestral nitrogenase sequences at the deepest nodes of the phylogeny suggest that ancient nitrogenases were Mo-dependent. We find that active-site sequence identity can reliably distinguish extant Mo-nitrogenases from V- and Fe-nitrogenases, as opposed to modeled active-site structural features that cannot be used to reliably classify nitrogenases of unknown metal dependence. Taxa represented by early-branching nitrogenase lineages lack one or more biosynthetic nifE and nifN genes that are necessary for assembly of the FeMo-cofactor, suggesting that early Mo-dependent nitrogenases may have utilized an alternate pathway for Mo-usage predating the FeMo-cofactor. Our results underscore the profound impacts that protein-level innovations likely had on shaping global biogeochemical cycles throughout Precambrian, in contrast to organism-level innovations which characterize Phanerozoic eon.
Massive sequencing of genetic markers, such as the 16S rRNA gene for prokaryotes, allows the comparative analysis of diversity and abundance of whole microbial communities. However, the data used for profiling microbial communities is usually low in signal and high in noise preventing the identification of real differences among treatments. PIME (Prevalence Interval for Microbiome Evaluation) fills this gap by removing those taxa that may be high in relative abundance in just a few samples but have a low prevalence overall. The reliability and robustness of PIME were compare against the existing methods and verified by a number of approaches using 16S rRNA independent datasets. To remove the noise, PIME filters microbial taxa not shared in a per treatment prevalence interval starting at 5% with increments of 5% at each filtering step. For each prevalence interval, hundreds of decision trees are calculated to predict the likelihood of detecting differences in treatments. The best prevalence-filtered dataset is user-selected by choosing the prevalence interval that keeps the majority of the 16S rRNA reads in the dataset and shows the lowest error rate. To obtain the likelihood of introducing bias while building prevalence-filtered datasets, an error detection step based in random permutations is also included. A reanalysis of previews published datasets with PIME uncovered previously missed microbial associations improving the ability to detect important organisms, which may be masked when only relative abundance is considered.
The data used for profiling microbial communities is usually sparse with some microbes having high abundance in a few samples and being nearly absent in others. However, current bioinformatics tools able to deal with this sparsity are lacking. pime (Prevalence Interval for Microbiome Evaluation) was designed to remove those taxa that may be high in relative abundance in just a few samples but have a low prevalence overall. The reliability and robustness of pime were compared against existing methods and tested using 16S rRNA independent data sets. pime filters microbial taxa not shared in a per treatment prevalence interval started at 5% prevalence with increasing increments of 5% at each filtering step. For each prevalence interval, hundreds of decision trees were calculated to predict the likelihood of detecting differences in treatments. The best prevalence-filtered data set was user-selected by choosing the prevalence interval that kept a large portion of the 16S rRNA sequences in the data set while also showing the lowest error rate. To obtain the likelihood of introducing type I error while building prevalence-filtered data sets, an error detection step based was also included. A pime reanalysis of published data sets uncovered other expected microbial associations than previously reported, which may be masked when only relative abundance was considered.
Prior to the identification of Xanthomonas perforans associated with bacterial spot of tomato in 1991, X. euvesicatoria was the only known species in Florida. Currently, X. perforans is the Xanthomonas sp. associated with tomato in Florida. Changes in pathogenic race and sequence alleles over time signify shifts in the dominant X. perforans genotype in Florida. We previously reported recombination of X. perforans strains with closely related Xanthomonas species as a potential driving factor for X. perforans evolution. However, the extent of recombination across the X. perforans genomes was unknown. We used a core genome multilocus sequence analysis approach to identify conserved genes and evaluated recombination-associated evolution of these genes in X. perforans. A total of 1,356 genes were determined to be “core” genes conserved among the 58 X. perforans genomes used in the study. Our approach identified three genetic groups of X. perforans in Florida based on the principal component analysis (PCA) using core genes. Nucleotide variation in 241 genes defined these groups, that are referred as Phylogenetic-group Defining (PgD) genes. Furthermore, alleles of many of these PgD genes showed 100% sequence identity with X. euvesicatoria, suggesting that variation likely has been introduced by recombination at multiple locations throughout the bacterial chromosome. Site-specific recombinase genes along with plasmid mobilization and phage associated genes were observed at different frequencies in the three phylogenetic groups and were associated with clusters of recombinant genes. Our analysis of core genes revealed the extent, source, and mechanisms of recombination events that shaped the current population and genomic structure of X. perforans in Florida.
MicroRNAs (miRNAs) are approximately 22 nucleotide (nt) long and play important roles in post-transcriptional regulation in both plants and animals. In animals, precursor (pre-) miRNAs are ∼70 nt hairpins produced by Drosha cleavage of long primary (pri-) miRNAs in the nucleus. Exportin-5 (XPO5) transports pre-miRNAs into the cytoplasm for Dicer processing. Alternatively, pre-miRNAs containing a 5' 7-methylguanine (m7G-) cap can be generated independently of Drosha and XPO5. Here we identify a class of m7G-capped pre-miRNAs with 5' extensions up to 39 nt long. The 5'-extended pre-miRNAs are transported by Exportin-1 (XPO1). Unexpectedly, a long 5' extension does not block Dicer processing. Rather, Dicer directly cleaves 5'-extended pre-miRNAs by recognizing its 3' end to produce mature 3p miRNA and extended 5p miRNA both in vivo and in vitro. The recognition of 5'-extended pre-miRNAs by the Dicer Platform-PAZ-Connector (PPC) domain can be traced back to ancestral animal Dicers, suggesting that this previously unrecognized Dicer reaction mode is evolutionarily conserved. Our work reveals additional genetic sources for small regulatory RNAs and substantiates Dicer's essential role in RNAi-based gene regulation.
A feature of the physiological adaptation to spaceflight in Arabidopsis thaliana (Arabidopsis) is the induction of reactive oxygen species (ROS)-associated gene expression. The patterns of ROS-associated gene expression vary among Arabidopsis ecotypes, and the role of ROS signalling in spaceflight acclimation is unknown. What could differences in ROS gene regulation between ecotypes on orbit reveal about physiological adaptation to novel environments? Analyses of ecotype-dependent responses to spaceflight resulted in the elucidation of a previously uncharacterized gene (OMG1) as being ROS-associated. The OMG1 5 flanking region is an active promoter in cells where ROS activity is commonly observed, such as in pollen tubes, root hairs, and in other tissues upon wounding. qRT-PCR analyses revealed that upon wounding on Earth, OMG1 is an apparent transcriptional regulator of MYB77 and GRX480, which are associated with the ROS pathway. Fluorescence-based ROS assays show that OMG1 affects ROS production. Phylogenetic analysis of OMG1 and closely related homologs suggests that OMG1 is a distant, unrecognized member of the CONSTANS-Like protein family, a member that arose via gene duplication early in the angiosperm lineage and subsequently lost its first DNA-binding B-box1 domain. These data illustrate that members of the rapidly evolving COL protein family play a role in regulating ROS pathway functions, and their differential regulation on orbit suggests a role for ROS signalling in spaceflight physiological adaptation.
Ancestral protein sequence reconstruction is a powerful technique for explicitly testing hypotheses about the evolution of molecular function, allowing researchers to meticulously dissect how historical changes in protein sequence impacted functional repertoire by altering the protein's 3D structure. These techniques have provided concrete, experimentally validated insights into ancient evolutionary processes and help illuminate the complex relationship between protein sequence, structure, and function. Inferring the protein family phylogenies on which ancestral sequence reconstruction depends and reconstructing the sequences, themselves, are amenable to high-throughput computational analysis. However, determining the structures of ancestral-reconstructed proteins and characterizing their functions typically rely on time-consuming and expensive laboratory analyses, limiting most current studies to examining a relatively small number of specific hypotheses. For this reason, we have little detailed, unbiased information about how molecular function evolves across large protein family phylogenies. Here we describe a generalized protocol that integrates ancestral sequence reconstruction with structural homology modeling and structure-based molecular affinity prediction to characterize historical changes in protein function across families with thousands of individual sequences. We highlight key steps in the analysis protocol requiring particularly careful attention to avoid introducing potential errors as well as steps for which computationally efficient subroutines can be substituted for more intensive approaches, allowing researchers to scale the analysis up or down, depending on available resources and requirements for reproducibility and scientific rigor. In our view, this approach provides a compelling compliment to more laboratory-intensive procedures, generating important contextual information that can help guide detailed experiments.
Abstract Reconstruction of ancestral protein sequences using phylogenetic methods is a powerful technique for directly examining the evolution of molecular function. Although ancestral sequence reconstruction (ASR) is itself very efficient, downstream functional, and structural studies necessary to characterize when and how changes in molecular function occurred are often costly and time-consuming, currently limiting ASR studies to examining a relatively small number of discrete functional shifts. As a result, we have very little direct information about how molecular function evolves across large protein families. Here we develop an approach combining ASR with structure and function prediction to efficiently examine the evolution of ligand affinity across a large family of double-stranded RNA binding proteins (DRBs) spanning animals and plants. We find that the characteristic domain architecture of DRBs—consisting of 2–3 tandem double-stranded RNA binding motifs (dsrms)—arose independently in early animal and plant lineages. The affinity with which individual dsrms bind double-stranded RNA appears to have increased and decreased often across both animal and plant phylogenies, primarily through convergent structural mechanisms involving RNA-contact residues within the β1–β2 loop and a small region of α2. These studies provide some of the first direct information about how protein function evolves across large gene families and suggest that changes in molecular function may occur often and unassociated with major phylogenetic events, such as gene or domain duplications.
One goal of structural biology is to understand how a protein’s 3-dimensional conformation determines its capacity to interact with potential ligands. In the case of small chemical ligands, deconstructing a static protein-ligand complex into its constituent atom-atom interactions is typically sufficient to rapidly predict ligand affinity with high accuracy (>70% correlation between predicted and experimentally-determined affinity), a fact that is exploited to support structure-based drug design. We recently found that protein-DNA/RNA affinity can also be predicted with high accuracy using extensions of existing techniques, but protein-protein affinity could not be predicted with >60% correlation, even when the protein-protein complex was available.
Understanding the structural basis for evolutionary changes in protein function is central to molecular evolutionary biology and can help determine the extent to which functional convergence occurs through similar or different structural mechanisms. Here, we combine ancestral sequence reconstruction with functional characterization and structural modeling to directly examine the evolution of sequence-structure-function across the early differentiation of animal and plant Dicer/DCL proteins, which perform the first molecular step in RNA interference by identifying target RNAs and processing them into short interfering products. We found that ancestral Dicer/DCL proteins evolved similar increases in RNA target affinities as they diverged independently in animal and plant lineages. In both cases, increases in RNA target affinities were associated with sequence changes that anchored the RNA's 5'phosphate, but the structural bases for 5'phosphate recognition were different in animal versus plant lineages. These results highlight how molecular-functional evolutionary convergence can derive from the evolution of unique protein structures implementing similar biochemical mechanisms.