The early metabolism arising in a Thioester world gave rise to amino acids and their simple peptides. The catalytic activity of these early simple peptides became instrumental in the transition from Thioester World to a Phosphate World. This transition involved the appearances of sugar phosphates, nucleotides, and polynucleotides. The coupling of the amino acids and peptides to nucleotides and polynucleotides is the origin for the genetic code. Many of the key steps in this transition are seen in the catalytic cores of the nucleotidyltransferases, the class II tRNA synthetases (aaRSs) and the CCA adding enzyme. These catalytic cores are dominated by simple beta hairpin structures formed in the Thioester World. The code evolved from a proto-tRNA, a tetramer XCCA interacting with a proto-aminoacyl-tRNA synthetase (aaRS) activating Glycine and Proline. The initial expanded code is found in the acceptor arm of the tRNA, the operational code. It is the coevolution of the tRNA with the aaRSs that is at the heart of the origin and evolution of the genetic code. There is also a close relationship between the accretion models of the evolving tRNA and that of the ribosome.
Phosphate is essential for all living systems, serving as a building block of genetic and metabolic machinery. However, it is unclear how phosphate could have assumed these central roles on primordial Earth, given its poor geochemical accessibility. We used systems biology approaches to explore the alternative hypothesis that a protometabolism could have emerged prior to the incorporation of phosphate. Surprisingly, we identified a cryptic phosphate-independent core metabolism producible from simple prebiotic compounds. This network is predicted to support the biosynthesis of a broad category of key biomolecules. Its enrichment for enzymes utilizing iron-sulfur clusters, and the fact that thermodynamic bottlenecks are more readily overcome by thioester rather than phosphate couplings, suggest that this network may constitute a “metabolic fossil” of an early phosphate-free nonenzymatic biochemistry. Our results corroborate and expand previous proposals that a putative thioester-based metabolism could have predated the incorporation of phosphate and an RNA-based genetic system.PaperClip
Class II Aminoacyl-tRNA synthetases are a set of very ancient multi domain proteins. The evolution of the catalytic domain of Class II synthetases can be reconstructed from three peptidyl-hairpins. Further evolution from this primordial catalytic core leads to a split of the Class II synthetases into two divisions potentially associated with the operational code. The earliest form of this code likely coded predominantly Glycine (Gly), Proline (Pro), Alanine (Ala) and "Lysine"/Aspartic acid (Lys/Asp). There is a paradox in these synthetases beginning with a hairpin structure before the Genetic Code existed. A resolution is found in the suggestion that the primordial Aminoacyl synthetases formed in a transition from a Thioester world to a Phosphate ester world.
The evolution of the genetic code is mapped out starting with the aminoacyl tRNA-synthetases and their interaction with the operational code in the tRNA acceptor arm. Combining this operational code with a metric based on the biosynthesis of amino acids from the Citric acid, we come to the conclusion that the earliest genetic code was a Guanine Cytosine (GC) code. This has implications for the likely earliest positively charged amino acids. The progression from this pure GC code to the extant one is traced out in the evolution of the Large Ribosomal Subunit, LSU, and its proteins; in particular those associated with the Peptidyl Transfer Center (PTC) and the nascent peptide exit tunnel. This progression has implications for the earliest encoded peptides and their evolutionary progression into full complex proteins.
Recent reviews discussed the critical roles of apoptosis in human spermatogenesis and infertility. These reviews highlight the FasL-induced caspase cascade in apoptosis lending importance to our discovery of the pseudogene status of the Lfg5 gene in modern humans, Neanderthal and the Denisovan. This gene is a member of the ancient and highly conserved apoptosis Lifeguard family. This pseudogenization is the result of a premature stop codon at the 3'-end of exon 8 not found in any other ortholog. With the current exception of the domesticated bovine and buffalo, Lfg5's expression in mammals is testis-specific. A full analysis of this gene, its phylogenetic context and its recent hominin changes suggest its inactivation was likely under selection in human evolution.
Many platforms for genome-wide analysis of gene expression contain ‘redundant’ measures for the same gene. For example, the most highly utilized platforms for gene expression microarrays, Affymetrix GeneChip® arrays, have as many as ten or more probe sets for some genes. Occasionally, individual probe sets for the same gene report different trends in expression across experimental conditions, a situation that must be resolved in order to accurately interpret the data. We developed an algorithm, SCOREM, for determining the level of agreement between such probe sets, utilizing a statistical test of concordance, Kendall's W coefficient of concordance, and a graph-searching algorithm for the identification of concordant probe sets. We also present methods for consolidating concordant groups into a single value for its corresponding gene and for post hoc analysis of discordant groups. By combining statistical consolidation with sequence analysis, SCOREM possesses the unique ability to identify biologically meaningful discordant behaviors, including differing behaviors in alternate RNA isoforms and tissue-specific patterns of expression. When consolidating concordant behaviors, SCOREM outperforms other methods in detecting both differential expression and overrepresented functional categories.
Background Trichomonas vaginalis has an unusually large genome (∼160 Mb) encoding ∼60,000 proteins. With the goal of beginning to understand why some Trichomonas genes are present in so many copies, we characterized here a family of ∼123 Trichomonas genes that encode transmembrane adenylyl cyclases (TMACs). Methodology/Principal Findings The large family of TMACs genes is the result of recent duplications of a small set of ancestral genes that appear to be unique to trichomonads. Duplicated TMAC genes are not closely associated with repetitive elements, and duplications of flanking sequences are rare. However, there is evidence for TMAC gene replacements by homologous recombination. A high percentage of TMAC genes (∼46%) are pseudogenes, as they contain stop codons and/or frame shifts, or the genes are truncated. Numerous stop codons present in the genome project G3 strain are not present in orthologous genes of two other Trichomonas strains (S1 and B7RC2). Each TMAC is composed of a series of N-terminal transmembrane helices and a single C-terminal cyclase domain that has adenylyl cyclase activity. Multiple TMAC genes are transcribed by Trichomonas cloned by limiting dilution. Conclusions/Significance We conclude that one reason for the unusually large genome of Trichomonas is the presence of unstable families of genes such as those encoding TMACs that are undergoing massive gene duplication and concomitant development of pseudogenes.
This paper is an attempt to trace the evolution of the ribosome through the evolution of the universal P-loop GTPases that are involved with the ribosome in translation and with the attachment of the ribosome to the membrane. The GTPases involved in translation in Bacteria/Archaea are the elongation factors EFTu/EF1, the initiation factors IF2/aeIF5b + aeIF2, and the elongation factors EFG/EF2. All of these GTPases also contain the OB fold also found in the non GTPase IF1 involved in initiation. The GTPase involved in the signal recognition particle in most Bacteria and Archaea is SRP54.
Surfactants find wide commercial use as foaming agents, emulsifiers, and dispersants. Currently, surfactants are produced from petroleum, or from seed oils such as palm or coconut oil. Due to concerns with CO(2) emissions and the need to protect rainforests, there is a growing necessity to manufacture these chemicals using sustainable resources In this report, we describe the engineering of a native nonribosomal peptide synthetase pathway (i.e., surfactin synthetase), to generate a Bacillus strain that synthesizes a highly water-soluble acyl amino acid surfactant, rather than the water insoluble lipopeptide surfactin. This novel product has a lower CMC and higher water solubility than myristoyl glutamate, a commercial surfactant. This surfactant is produced by fermentation of cellulosic carbohydrate as feedstock. This method of surfactant production provides an approach to sustainable manufacturing of new surfactants.
Numerous protists and rare fungi have truncated Asn-linked glycan precursors and lack N-glycan-dependent quality control (QC) systems for glycoprotein folding in the endoplasmic reticulum. Here, we show that the abundance of sequons (NXT or NXS), which are sites for N-glycosylation of secreted and membrane proteins, varies by more than a factor of 4 among phylogenetically diverse eukaryotes, based on a few variables. There is positive correlation between the density of sequons and the AT content of coding regions, although no causality can be inferred. In contrast, there appears to be Darwinian selection for sequons containing Thr, but not Ser, in eukaryotes that have N-glycan-dependent QC systems. Selection for sequons with Thr, which nearly doubles the sequon density in human secreted and membrane proteins, occurs by an increased conditional probability that Asn and Thr are present in sequons rather than elsewhere. Increasing sequon densities of the hemagglutinin (HA) of influenza viruses A/H3N2 and A/H1N1 during the past few decades of human infection also result from an increased conditional probability that Asn, Thr, and Ser are present in sequons rather than elsewhere. In contrast, there is no selection on sequons by this mechanism in HA of A/H5N1 or 2009 A/H1N1 (Swine flu). Very strong selection for sequons with both Thr and Ser in glycoprotein of M-r 120,000 (gp120) of HIV and related retroviruses results from this same mechanism, as well as amino acid composition bias and increases in AT content. We conclude that there is Darwinian selection for sequons in phylogenetically disparate eukaryotes and viruses.
The expanding wealth of human, model and other organism’s genomic data has allowed the identification of a distinct gene family of apoptotic related genes. Most of these genes are currently unannotated or have been subsumed under two questionably related gene families in the past. For example the transmembrane Bax inhibitor 1 (BI1) motif family has been reported to play a role in apoptosis and to consist of at least seven mammalian protein genes, GRINA , BI1 , Lfg / FAIM2 , Ghitm , RESC1 / Tmbim1 , GAAP / Tmbim4 , and Tmbm1b . However, a detailed sequence and phylogenetic analysis shows that only five of these form a clear and unique protein family. This now provides information for understanding and investigating the biological roles of these proteins across a wide range of tissues in model organisms. The evolutionary relationships among these genes provide a powerful prospective for extrapolating to human conditions.
The Cilium, the Nucleus and the Mitochondrion are three important organelles whose evolutionary histories are intimately related to the evolution and origin of the eukaryotic cell. The cilium is involved in motility and sensory transduction. The cilium is only found in the eukaryotic cells. Here we show that eight gene duplications prior to the last common ancestor of all extant eukaryotes account for the expansion of the Heavy Chain Dynein family of motor proteins and the evolution of the complexity of the cilium. The ambiguities in the branching of the phylogenetic tree of the HC-Dyneins were resolved by creating well-defined subtrees and using them to create the full tree. Due to the intimate relationship between the nucleus, the division center, mitosis and the basal body/centriole, the evolution of the cilium can now be related to the evolution of mitosis. In addition, the analysis of the cilium rules out its endosymbiotic origin from a phagocytosis of a bacterium.
Fractures are among the most common human traumas. Fracture healing represents a unique temporarily definable post-natal process in which to study the complex interactions of multiple molecular events that regulate endochondral skeletal tissue formation. Because of the regenerative nature of fracture healing, it is hypothesized that large numbers of post-natal stem cells are recruited and contribute to formation of the multiple cell lineages that contribute to this process. Bayesian modeling was used to generate the temporal profiles of the transcriptome during fracture healing. The temporal relationships between ontologies that are associated with various biologic, metabolic, and regulatory pathways were identified and related to developmental processes associated with skeletogenesis, vasculogenesis, and neurogenesis. The complement of all the expressed BMPs, Wnts, FGFs, and their receptors were related to the subsets of transcription factors that were concurrently expressed during fracture healing. We further defined during fracture healing the temporal patterns of expression for 174 of the 193 genes known to be associated with human genetic skeletal disorders. In order to identify the common regulatory features that might be present in stem cells that are recruited during fracture healing to other types of stem cells, we queried the transcriptome of fracture healing against that seen in embryonic stem cells (ESCs) and mesenchymal stem cells (MSCs). Approximately 300 known genes that are preferentially expressed in ESCs and ∼350 of the known genes that are preferentially expressed in MSCs showed induction during fracture healing. Nanog, one of the central epigenetic regulators associated with ESC stem cell maintenance, was shown to be associated in multiple forms or bone repair as well as MSC differentiation. In summary, these data present the first temporal analysis of the transcriptome of an endochondral bone formation process that takes place during fracture healing. They show that neurogenesis as well as vasculogenesis are predominant components of skeletal tissue formation and suggest common pathways are shared between post-natal stem cells and those seen in ESCs.
Comparative analysis of closely related genomes is expected to yield significant insights into the processes of evolution, development, and regulation. The recently sequenced genomes of twelve fruit fly (genus Drosophila ) species and other insects provide an ideal data set for this purpose. The primary focus of this research effort is on computational analysis of genome synteny (analysis of relative gene-order conservation between species), chromosomal dynamics, and genome rearrangement between species. We have developed computational methods to process draft genome assemblies and to infer cross-species synteny. Additionally, we have developed computationally efficient algorithms to infer evolutionary rearrangement event counts and ancestral synteny blocks with the ability to handle a large set of species with high "gene counts". Finally, we analyzed chromosomal rearrangements due to large-scale events such as multi-gene inversions, and fine-scale events such as single-gene relocation. Our work provides new methodologies to enable fast comparative analysis of multi-species genome-scale datasets. Our results open a window into evolutionary chromosomal reorganization within a set of eukaryotic species and highlight the role of large-scale and fine-scale rearrangement events.
The sequencing of the 12 genomes of members of the genus Drosophila was taken as an opportunity to reevaluate the genetic and physical maps for 11 of the species, in part to aid in the mapping of assembled scaffolds. Here, we present an overview of the importance of cytogenetic maps to Drosophila biology and to the concepts of chromosomal evolution. Physical and genetic markers were used to anchor the genome assembly scaffolds to the polytene chromosomal maps for each species. In addition, a computational approach was used to anchor smaller scaffolds on the basis of the analysis of syntenic blocks. We present the chromosomal map data from each of the 11 sequenced non-Drosophila melanogaster species as a series of sections. Each section reviews the history of the polytene chromosome maps for each species, presents the new polytene chromosome maps, and anchors the genomic scaffolds to the cytological maps using genetic and physical markers. The mapping data agree with Muller's idea that the majority of Drosophila genes are syntenic. Despite the conservation of genes within homologous chromosome arms across species, the karyotypes of these species have changed through the fusion of chromosomal arms followed by subsequent rearrangement events.
Background: The origin and early evolution of the active site of the ribosome can be elucidated through an analysis of the ribosomal proteins' taxonomic block structures and their RNA interactions. Comparison between the two subunits, exploiting the detailed three-dimensional structures of the bacterial and archaeal ribosomes, is especially informative.Results: The analysis of the differences between these two sites can be summarized as follows: 1) There is no self-folding RNA segment that defines the decoding site of the small subunit; 2) there is one self-folding RNA segment encompassing the entire peptidyl transfer center of the large subunit; 3) the protein contacts with the decoding site are made by a set of universal alignable sequence blocks of the ribosomal proteins; 4) the majority of those peptides contacting the peptidyl transfer center are made by bacterial or archaeal-specific sequence blocks.Conclusion: These clear distinctions between the two subunit active sites support an earlier origin for the large subunit's peptidyl transferase center (PTC) with the decoding site of the small subunit being a later addition to the ribosome. The main implications are that a single self-folding RNA, in conjunction with a few short stabilizing peptides, formed the precursor of the modern ribosomal large subunit in association with a membrane.Reviewers: This article was reviewed by Jerzy Jurka, W. Ford Doolittle, Eugene Shaknovich, and George E. Fox (nominated by Jerzy Jurka).
The recognition of the role of mathematics and computer science in modern biology has led to new terminology, as did chemistry with biochemistry, and physics with biophysics. We need to think only of bioinformatics, computational biology, and even system biology and genomics for example. These terms seem to strongly suggest that this is all rather new. Yet a short review of the work of those such as J.B.S. Haldane, Sewell Wright, DArcy Thompson and R.A. Fisher, to say nothing of scientists like Luria and Delbrueck or Hodgkin and Huxley or Thomas Hunt Morgan, is useful. Their work and foresight set the stage for modern applications of mathematical modeling and statistics in the biological sciences. It has often been said that the only difference between now and then is the increase in data—a lot more data. This is clearly not the full story. In addition, we have computational power unimaginable to these earlier researchers, as well as to anyone only forty years ago. So what are our challenges? Some are clear, including the modeling and analysis of biologys complex systems such as a cells signaling, metabolic and differentiation. Also needed are analysis and models of complex neural systems and ecological structures. The latter, for example, will require a nearly full revamping of the early field of population genetics and evolution in order to exploit both modern genomics and new field studies of multiple species and environmental interactions. And there will be more, much of which will only become apparent as new data and questions arise. One example would be RNAi and micro-arrays inducing the development of new analysis tools. About the keynote speaker. Dr. Temple Smith graduated with a Ph.D. in Nuclear Physics from University of Colorado. He did a joint postdoctoral fellowship under the direction of the mathematician, Stanislaw Ulam and the molecular biologist, John Sadler. He was one of the founders of GenBank at Los Alamos. Dr. Smith has been the Director of the BioMolecular Engineering Research Center in the College of Engineering at Boston University since 1991. He is a professor in the Department of Biomedical Engineering and co-founder of the company, Modular Genetics, Inc. Dr. Smith is a co-developer of the Smith-Waterman sequence alignment algorithm, the standard tool used in most DNA and protein sequence comparison. His research is centered on the application of various computer science and mathematical methods to the discovery of the syntactic and semantic patterns in nucleic acid and amino acid sequences. These include the development of new sequence pattern extraction tools, multidomain dissection methods, and protein inverse folding prediction algorithms. In addition, Dr. Smith has carried out research in the application of many such methods ranging from the time calibration of HIV viral evolution analysis and modeling of the WD repeat family of proteins, to ribosomal protein evolution.
Richard Lathrop合作论文数University of California, Irvine;Information and Computer Science Department3