BACKGROUND:Molecular evolution is a very active field of research, with several complementary approaches, including dN/dS, HON90, MM01, and others. Each has documented strengths and weaknesses, and no one approach provides a clear picture of how natural selection works at the molecular level. The purpose of this work is to present a simple new method that uses quantitative amino acid properties to identify and characterize directional selection in proteins.METHODS:Inferred amino acid replacements are viewed through the prism of a single physicochemical property to determine the amount and direction of change caused by each replacement. This allows the calculation of the probability that the mean change in the single property associated with the amino acid replacements is equal to zero (H0: μ = 0; i.e., no net change) using a simple two-tailed t-test.RESULTS:Example data from calanoid and cyclopoid copepod cytochrome oxidase subunit I sequence pairs are presented to demonstrate how directional selection may be linked to major shifts in adaptive zones, and that convergent evolution at the whole organism level may be the result of convergent protein adaptations.CONCLUSIONS:Rather than replace previous methods, this new method further complements existing methods to provide a holistic glimpse of how natural selection shapes protein structure and function over evolutionary time.
Amino acid property methods have repeatedly proven more sensitive than nucleotide-based methods, especially to the subtle influences of selection on Single Nucleotide Polymorphisms (SNPs). The purpose of this study is to evaluate the effects of population sampling and sliding window size on the sensitivity of one amino acid based method by assessing drug resistant SNPs from the HIV-1 gag-pol gene. The analysis of most of the SNPs produced positive results. There was not a trend in terms of properties affected most often, but sliding window size affected results more than population sampling.
TreeSAAP has been used in a variety of protein studies for detecting adaptation in terms of the physicochemical properties involved in amino acid replacement. The accuracy of TreeSAAP was here tested using simulated protein-coding DNA data. A sampling of 1402 simulated amino acid replacements resulted in a default accuracy of 81.1%, with most properties exhibiting >90% accuracy. More than half of the false-positive results were traced to just 11 of the 180 possible single-step amino acid exchanges. Overall accuracy increased as the number of magnitude partitions used in the analysis decreased. Sliding window size did not significantly affect accuracy.
BACKGROUND:Metabolism of energy nutrients by the mitochondrial electron transport chain (ETC) is implicated in the aging process. Polymorphisms in core ETC proteins may have an effect on longevity. Here we investigate the cytochrome b (cytb) polymorphism at amino acid 7 (cytbI7T) that distinguishes human mitochondrial haplogroup H from haplogroup U.PRINCIPAL FINDINGS:We compared longevity of individuals in these two haplogroups during historical extremes of caloric intake. Haplogroup H exhibits significantly increased longevity during historical caloric restriction compared to haplogroup U (p = 0.02) while during caloric abundance they are not different. The historical effects of natural selection on the cytb protein were estimated with the software TreeSAAP using a phylogenetic reconstruction for 107 mammal taxa from all major mammalian lineages using 13 complete protein-coding mitochondrial gene sequences. With this framework, we compared the biochemical shifts produced by cytbI7T with historical evolutionary pressure on and near this polymorphic site throughout mammalian evolution to characterize the role cytbI7T had on the ETC during times of restricted caloric intake.SIGNIFICANCE:Our results suggest the relationship between caloric restriction and increased longevity in human mitochondrial haplogroup H is determined by cytbI7T which likely enhances the ability of water to replenish the Q(i) binding site and decreases the time ubisemiquinone is at the Q(o) site, resulting in a decrease in the average production rate of radical oxygen species (ROS).
We present a new algorithm, ChemAlign, that uses physicochemical properties and secondary structure elements to create biologically relevant multiple sequence alignments (MSAs). Additionally, we introduce the physicochemical property difference (PPD) score for the evaluation of MSAs. This score is the normalized difference of physicochemical property values between a calculated and a reference alignment. It takes a step beyond sequence similarity and measures characteristics of the amino acids to provide a more biologically relevant metric. ChemAlign is able to produce more biologically correct alignments and can help to identify potential drug docking sites.
Quinoa (Chenopodium quinoa Willd.) is a food crop cultivated by subsistence farmers and commercial growers on the high Andean plateau, primarily in Bolivia, Peru, and Chile. Present interest in quinoa is due to its tolerance of harsh environments and its nutritional value. It is thought that the seed storage proteins of quinoa, particularly the 11S globulins and 2S albumins, are responsible for the relatively high protein content and ideal amino acid balance of the quinoa seed. Here we report the genomic and cDNA sequences for two 11S genes representing two orthologous loci from the quinoa genome. Important features of the genes and the proteins they encode are described on the basis of a comparison with homologous 11S sequences from other plant species. Gene expression and protein accumulation determined via reverse transcriptase real‐time PCR and SDS‐PAGE analyses are described. Additionally, we report the phylogenetic relationships between quinoa and 49 other species by using the coding DNA sequence for the well‐conserved 11S basic subunit.
Foot-and-mouth disease virus is an economically important animal virus that exhibits extensive genetic and antigenic heterogeneity. To examine the evolutionary forces that have influenced the population dynamics of foot-and-mouth disease virus, individual genes and the coding genomes for the Eurasian (Asia1, A, C, and O) serotypes were examined for phylogenetic relationships, recombination, genetic diversity and selection. Our analyses demonstrate that paraphyletic relationships among serotypes are not as prevalent as previously proposed and suggest that convergent evolution might be obscuring phylogenetic relationships. We provide evidence that identification of recombinant sequences and recombination breakpoint patterns among and within serotypes are heavily dependent on the level of genetic diversity and convergent characters present in a particular data set as well as the methods used to detect recombination. Here, we also investigate the impact of adaptive positive selection on the capsid proteins and the non-structural genes 2B, 2C, 3A, and 3Cpro to identify genome regions involved in genetic diversity and antigenic variation. Two different categories of positive selection at the amino acid level were examined; conservative (stabilizing) selection that maintains particular phenotypic properties of an amino acid residue and radical (destabilizing), and selection that dramatically alters the phenotype and potentially the functional and/or structural features of the protein. Approximately, 29% of residues in the capsid proteins were under positive selection. Of those, 64% were under the influence of destabilizing selection, 80% were under the influence of stabilizing selection, and 44% had phenotypic properties influenced by both selection types. The majority of residues under selection (74%) were located outside of known antigenic sites; suggestive of additional uncharacterized epitopes and genomic regions involved in antigenic drift.
Cytochrome c(1) (cyt-c(1)) and the Rieske Iron Sulphur Protein (ISP) are subunits of the cytochrome bc(1) complex located in the mitochondria functioning both as a proton pump and an electron transporter. Vertebrate model organism phylogenies were used in conjunction with existing 3D protein structures to evaluate the biochemical evolution of cyt-c(1) and ISP in terms of selection on amino acid properties. We found selection acting on the exterior surfaces of both proteins and specifically the core region of cyt-c(1). There is evidence supporting coevolution of these proteins relative to alpha helical tendencies, compressibility and equilibrium constant.
The CYP2D6 gene is responsible for metabolising a large portion of the commonly prescribed drugs. Because of its importance, various approaches have been taken to analyse CYP2D6 and Single Nucleotide Polymorphisms (SNPs) throughout its sequence. This study introduces a novel method to analyse the effects of SNPs on encoded protein complexes by focusing on the biochemical properties of each non-synonymous substitution using the program TreeSAAP. Our results show four SNPs in CYP2D6 that exhibit radical changes in amino acid properties which may cause a lack of functionality in the CYP2D6 gene and contribute to a person's inability to metabolise specific drugs.
We investigated whether the effect of evolutionary selection on three recent Single Nucleotide Polymorphisms (SNPs) in the mitochondrial sub-haplogroups of Pima Indians is consistent with their effects on metabolic efficiency. The mitochondrial SNPs impact metabolic rate and respiratory quotient, and may be adaptations to caloric restriction in a desert habitat. Using TreeSAAP software, we examined evolutionary selection in 107 mammalian species at these SNPs, characterising the biochemical shifts produced by the amino acid substitutions. Our results suggest that two SNPs were affected by selection during mammalian evolution in a manner consistent with their effects on metabolic efficiency in Pima Indians.
MOTIVATIONMultiple sequence alignments (MSAs) are at the heart of bioinformatics analysis. Recently, a number of multiple protein sequence alignment benchmarks (i.e. BAliBASE, OXBench, PREFAB and SMART) have been released to evaluate new and existing MSA applications. These databases have been well received by researchers and help to quantitatively evaluate MSA programs on protein sequences. Unfortunately, analogous DNA benchmarks are not available, making evaluation of MSA programs difficult for DNA sequences.RESULTSThis work presents the first known multiple DNA sequence alignment benchmarks that are (1) comprised of protein-coding portions of DNA (2) based on biological features such as the tertiary structure of encoded proteins. These reference DNA databases contain a total of 3545 alignments, comprising of 68 581 sequences. Two versions of the database are available: mdsa_100s and mdsa_all. The mdsa_100s version contains the alignments of the data sets that TBLASTN found 100% sequence identity for each sequence. The mdsa_all version includes all hits with an E-value score above the threshold of 0.001. A primary use of these databases is to benchmark the performance of MSA applications on DNA data sets. The first such case study is included in the Supplementary Material.
1 History 1 Summary 1 2 2 . DISCOVERY AND EPIDEMIOLOGY OF THE HUMAN POLYOMAVIRUSES BK VIRUS (BKV) AND JC VIRUS (JCV) .. . 1 9 Wendy A. Knowles Abstract 1 9 Discovery of BKV and JCV 1 9 Epidemiology 201 9 Discovery of BKV and JCV 1 9 Epidemiology 20 3. PHYLOGENOMICS AND MOLECULAR EVOLUTIO N OF POLYOMAVIRUSES 4 6 Keith A . Crandall, Marcos P rez-Losada, Ryan G . Christensen, David A. McClellan and Raphael P. Viscidi Abstract 46 Introduction 46 Population Variation of SV40 56 Summary 5646 Introduction 46 Population Variation of SV40 56 Summary 56 4. VIRUS RECEPTORS AND TROPISM 6 0 Aarthi Ashok and Walter J. Atwood Abstract 60 Introduction 60 Mouse Polyomavirus (PyV) 61 JC Virus (JCV) 63 B-Lymphotropic Papovavirus (LPV) 66 Conclusions 6860 Introduction 60 Mouse Polyomavirus (PyV) 61 JC Virus (JCV) 63 B-Lymphotropic Papovavirus (LPV) 66 Conclusions 68 5 . SEROLOGICAL CROSS REACTIVITY BETWEE N POLYOMAVIRUS CAPSIDS 73 Raphael P. Viscidi and Barbara Clayman Abstract 73 Introduction 73 Virus-Like Particle-Based Polyomavirus Enzyme Immunoassays 74 Reactivity of Rhesus Macaque Sera in VLP-Based Polyomaviru s Enzyme Immunoassays 76 Reactivity of Human Sera in VLP-Based Polyomavirus Enzyme Immunoassays 78 Reactivity of Human Sera in LPV VLP-Based Enzyme Immunoassay 8 1 Conclusions 8 173 Introduction 73 Virus-Like Particle-Based Polyomavirus Enzyme Immunoassays 74 Reactivity of Rhesus Macaque Sera in VLP-Based Polyomaviru s Enzyme Immunoassays 76 Reactivity of Human Sera in VLP-Based Polyomavirus Enzyme Immunoassays 78 Reactivity of Human Sera in LPV VLP-Based Enzyme Immunoassay 8 1 Conclusions 8 1 6. MOLECULAR GENETICS OF THE BK VIRUS 8 5 Christopher L . Cubitt Abstract 85 Introduction 85 Genotyping BKV 90 Association of Mutations and Genotypes with Disease 90 BKV Molecular Genetics: Future Research 9285 Introduction 85 Genotyping BKV 90 Association of Mutations and Genotypes with Disease 90 BKV Molecular Genetics: Future Research 92 7. SEROLOGICAL DIAGNOSIS OF HUMAN POLYOMAVIRU S INFECTION 9 6 Annika Lundstig and Joakim Dillne r Abstract 96 Introduction 96 Simian Virus 40 97 Antibody Stability 98 Serological Methods 98 EIA Serological Method 99 Preparation of Virus-Like Particles 99 Expression Systems 10096 Introduction 96 Simian Virus 40 97 Antibody Stability 98 Serological Methods 98 EIA Serological Method 99 Preparation of Virus-Like Particles 99 Expression Systems 100 8. HUMAN POLYOMAVIRUS JC AND BK PERSISTEN T INFECTION 10 2 Kristina Doerries Abstract 102 Introduction 102 Urogenital Persistent BKV Infection 105 Active BKV Infection in the Kidney 105 JCV in the Urogenital Tract 106 Activation of Urogenital JCV Infection 107 Neurotropism of Human Polyomaviruses 108 JCV in the Central Nervous System 108 Activity of Asymptomatic JCV Infection in the CNS 11 0 Human Polyomaviruses in the Hematopoietic System 111102 Introduction 102 Urogenital Persistent BKV Infection 105 Active BKV Infection in the Kidney 105 JCV in the Urogenital Tract 106 Activation of Urogenital JCV Infection 107 Neurotropism of Human Polyomaviruses 108 JCV in the Central Nervous System 108 Activity of Asymptomatic JCV Infection in the CNS 11 0 Human Polyomaviruses in the Hematopoietic System 111 Association of BKV with Cells of the Immune System 11 2 JCV in Lymphoid Organs and Blood Cells 11 3 Activation of JCV Infection in Hematopoietic Cells 11 4 9 . IMMUNITY AND AUTOIMMUNITY INDUCED B Y POLYOMAVIRUSES : CLINICAL, EXPERIMENTA L AND THEORETICAL ASPECTS 11 7 Ole Petter Rekvig, Signy Bendiksen and Ugo Moen s Abstract 11 7 Introduction 11 7 Immunology of Polyomaviruses 11 9 Polyomaviruses and T Cell Responses 12 6 Polyomaviruses, SLE and Autoimmunity to Nucleosomes and dsDNA 13 0 Concluding Remarks 13 911 7 Introduction 11 7 Immunology of Polyomaviruses 11 9 Polyomaviruses and T Cell Responses 12 6 Polyomaviruses, SLE and Autoimmunity to Nucleosomes and dsDNA 13 0 Concluding Remarks 13 9 10. THE PATHOBIOLOGY OF POLYOMAVIRUS INFECTION IN MAN 148 Parmjeet Randhawa, Abhay Vats and Ron Shapiro Abstract 14 8 Biology of Polyomaviruses 14 8 Historical Aspects 149 Modes of Natural Transmission 149 Transmission of Polyomavirus via Organ Transplantation 150 Viral Interactions with Host Cell Receptors 150 Entry of Virus into Host Cells 15 1 Cytoplasmic Trafficking 152 Nuclear Targeting 153 Clinical Sequelae of Primary Infection in Man 1153 Sites of Viral Latency 153 Reactivation of Latent Virus 154 Pathogenesis of Tissue Damage in Polyomavirus Infected Tissues 154 Concluding Remarks » 15514 8 Biology of Polyomaviruses 14 8 Historical Aspects 149 Modes of Natural Transmission 149 Transmission of Polyomavirus via Organ Transplantation 150 Viral Interactions with Host Cell Receptors 150 Entry of Virus into Host Cells 15 1 Cytoplasmic Trafficking 152 Nuclear Targeting 153 Clinical Sequelae of Primary Infection in Man 1153 Sites of Viral Latency 153 Reactivation of Latent Virus 154 Pathogenesis of Tissue Damage in Polyomavirus Infected Tissues 154 Concluding Remarks » 155 11 . POLYOMAVIRUS-ASSOCIATED NEPHROPATHY IN RENAL TRANSPLANTATION: CRITICAL ISSUES OF SCREENING AND MANAGEMENT 16 0 Hans H. Hirsch, Cinthia B. Drachenberg, Juerg Steiger and Emilio Ramos Abstract 160 Introduction 160 Infection, Replication and Disease 161 Epidemiology 163 Risk Factors 164 Diagnosis 166 Intervention 1169 Retransplantation 170 Conclusion 170160 Introduction 160 Infection, Replication and Disease 161 Epidemiology 163 Risk Factors 164 Diagnosis 166 Intervention 1169 Retransplantation 170 Conclusion 170 12. BK VIRUS AND IMMUNOSUPPRESSIVE AGENTS 174 Irfan Agha and Daniel C . Brennan Abstract 174 Biology of BKV Infection : From Primary Infection to Manifest Disease 174 The Second Hit Hypothesis 175 Clinical Correlates of BKV Infection in Renal Transplant Recipients 176 Impact of Immunosuppression : Net State or Asymmetric Predisposition? 176 Prospective Look at BKVN : Role of Immunosuppression 178 Reflections on the Biologic Behavior of the Virus 180 Risks of Alteration of Immunosuppression 181 Immunosuppression Management for BKV Infection 181174 Biology of BKV Infection : From Primary Infection to Manifest Disease 174 The Second Hit Hypothesis 175 Clinical Correlates of BKV Infection in Renal Transplant Recipients 176 Impact of Immunosuppression : Net State or Asymmetric Predisposition? 176 Prospective Look at BKVN : Role of Immunosuppression 178 Reflections on the Biologic Behavior of the Virus 180 Risks of Alteration of Immunosuppression 181 Immunosuppression Management for BKV Infection 181 13 . BK VIRUS INFECTION AFTER NON-RENAL TRANSPLANTATION 185 Martha Pavlakis, Abdolreza Haririan and David K . Klassen Abstract 185 BKV-BMT-Hemorrhagic Cystitis 185 BKV-Non-Renal Solid Organ Transplant 188185 BKV-BMT-Hemorrhagic Cystitis 185 BKV-Non-Renal Solid Organ Transplant 188 14. LATENT AND PRODUCTIVE POLYOMAVIRUS INFECTION S OF RENAL ALLOGRAFTS: MORPHOLOGICAL , CLINICAL, AND PATHOPHYSIOLOGICAL ASPECTS 190 Volker Nickeleit, Harsharan K. Singh and Michael J . Mihatsch Abstract 19 0 Introduction 19 0 Morphologic Characterization of BK Virus Allograft Nephropathy (BKN) 19 1 Ultrastructural Features 193 Ancillary Diagnostic Techniques 193 Histologic Stages/Patterns of BKN 196 Latent BK Virus Infections 19819 0 Introduction 19 0 Morphologic Characterization of BK Virus Allograft Nephropathy (BKN) 19 1 Ultrastructural Features 193 Ancillary Diagnostic Techniques 193 Histologic Stages/Patterns of BKN 196 Latent BK Virus Infections 198 15. URINE CYTOLOGY FINDINGS OF POLYOMAVIRU S INFECTIONS 20 1 Harsharan K. Singh, Lukas Bubendorf, Michael J . Mihatsch , Cinthia B. Drachenberg and Volker Nickelei t Abstract 20
We use a multigene data set (the mitochondrial locus and nine nuclear gene regions) to test phylogenetic relationships in the South American "lava lizards" (genus Microlophus) and describe a strategy for aligning noncoding sequences that accounts for differences in tempo and class of mutational events. We focus on seven nuclear introns that vary in size and frequency of multibase length mutations (i.e., indels) and present a manual alignment strategy that incorporates insertions and deletions (indels) for each intron. Our method is based on mechanistic explanations of intron evolution that does not require a guide tree. We also use a progressive alignment algorithm (Probabilistic Alignment Kit; PRANK) and distinguishes insertions from deletions and avoids the "gapcost" conundrum. We describe an approach to selecting a guide tree purged of ambiguously aligned regions and use this to refine PRANK performance. We show that although manual alignment is successful in finding repeat motifs and the most obvious indels, some regions can only be subjectively aligned, and there are limits to the size and complexity of a data matrix for which this approach can be taken. PRANK alignments identified more parsimony-informative indels while simultaneously increasing nucleotide identity in conserved sequence blocks flanking the indel regions. When comparing manual and PRANK with two widely used methods (CLUSTAL, MUSCLE) for the alignment of the most length-variable intron, only PRANK recovered a tree congruent at deeper nodes with the combined data tree inferred from all nuclear gene regions. We take this concordance as an objective function of alignment quality and present a strongly supported phylogenetic hypothesis for Microlophus relationships. From this hypothesis we show that (1) a coded indel data partition derived from the PRANK alignment contributed significantly to nodal support and (2) the indel data set permitted detection of significant conflict between mitochondrial and nuclear data partitions, which we hypothesize arose from secondary contact of distantly related taxa, followed by hybridization and mtDNA introgression.
Cytochrome c(1) (cyt-c(1)) and the Rieske Iron Sulphur Protein (ISP) are subunits of the cytochrome bc(1) complex located in the mitochondria functioning both as a proton pump and an electron transporter. Vertebrate model organism phylogenies were used in conjunction with existing 3D protein structures to evaluate the biochemical evolution of cyt-c(1) and ISP in terms of selection on amino acid properties. We found selection acting on the exterior surfaces of both proteins and specifically the core region of cyt-c(1). There is evidence supporting coevolution of these proteins relative to alpha helical tendencies, compressibility and equilibrium constant.
We provide in this chapter an overview of the basic steps to reconstruct evolutionary relationships through standard phylogeny estimation approaches as well as network approaches for sequences more closely related. We discuss the importance of sequence alignment, selecting models of evolution, and confidence assessment in phylogenetic inference. We also introduce the reader to a variety of software packages used for such studies. Finally, we demonstrate these approaches throughout using a data set of 33 whole genomes of polyomaviruses. A robust phylogeny of these genomes is estimated and phylogenetic relationships among the polyomaviruses determined using Bayesian and maximum likelihood approaches. Furthermore, population samples of SV40 are used to demonstrate the utility of network approaches for closely related sequences. The phylogenetic analysis suggested a close relationship among the BK viruses, JC viruses, and SV40 with a more distant association with mouse polyomavirus, monkey polymavirus (LPV) and then avian polyomavirus (BFDV).
Investigations of opsin evolution outside of vertebrate systems have long been focused on insect visual pigments, whereas other groups have received little attention. Furthermore, few studies have explicitly investigated the selective influences across all the currently characterized arthropod opsins. In this study, we contribute to the knowledge of crustacean opsins by sequencing 1 opsin gene each from 6 previously uncharacterized crustacean species (Euphausia superba, Homarus gammarus, Archaeomysis grebnitzkii, Holmesimysis costata, Mysis diluviana, and Neomysis americana). Visual pigment spectral absorbances were measured using microspectrophotometry for species not previously characterized (A. grebnitzkii=496 nm, H. costata=512 nm, M. diluviana=501 nm, and N. americana=520 nm). These novel crustacean opsin sequences were included in a phylogenetic analysis with previously characterized arthropod opsin sequences to determine the evolutionary placement relative to the well-established insect spectral clades (long-/middle-/short-wavelength sensitive). Phylogenetic analyses indicate these novel crustacean opsins form a monophyletic clade with previously characterized crayfish opsin sequences and form a sister group to insect middle-/long-wavelength-sensitive opsins. The reconstructed opsin phylogeny and the corresponding spectral data for each sequence were used to investigate selective influences within arthropod, and mainly "pancrustacean," opsin evolution using standard dN/dS ratio methods and more sensitive techniques investigating the amino acid property changes resulting from nonsynonymous replacements in a historical (i.e., phylogenetic) context. Although the conservative dN/dS methods did not detect any selection, 4 amino acid properties (coil tendencies, compressibility, power to be at the middle of an alpha-helix, and refractive index) were found to be influenced by destabilizing positive selection. Ten amino acid sites relating to these properties were found to face the binding pocket, within 4 A of the chromophore and thus have the potential to affect spectral tuning.
The nicotinic acetylcholine receptor (nAChR) is a well characterized ion channel, but studies have not yet resolved the effects of positive selection on individual physicochemical amino acid properties, which could be used to identify protein structure and function a priori. An analytical software package, TreeSAAP, was used to detect selection and molecular adaptation based on evolutionary pathways and amino acid substitutions that produced changes in physicochemical properties in the alpha subunit of nAChR. These studies were coupled with visualization software, allowing individual codons under selective pressure to be viewed in their proper secondary, tertiary, and quaternary structural context. We present here the sites affected by positive destabilizing selection in nAChR and characterize their respective physicochemical property shifts. The data suggest that during nAChR evolution, specific properties were preferentially selected for at residues lining the aqueous interfaces, subunit interaction regions and at regions known to interact with various recruitment enzymes. The nicotinic acetylcholine receptor (nAChR) is a protein within the nervous system responsible for synaptic initiation of muscle contraction at the neuromuscular junction (NMJ). It has been implicated in several neurological disorders including Parkinson’s, Alzheimer’s, and a rare condition known as congenital myasthenic syndrome. Being the best characterized of the 4-pass transmembrane helix family of ion channels, studies have been done to resolve structural (Miyazawa et al. 2003), functional (Unwin et al. 2003), and phylogenetic characteristics (Novère et al., 1995), as well as some subunit and protein-protein interactions (see Discussion). However, the implications of phylogenetic effects on the amino acid level have yet to be resolved. In order to identify the effect of positive destabilizing selection (McClellan et al., 2005) amino acid substitutions and changes in respective properties must be established and analyzed. This is done using TreeSAAP (Wooley et al. 2003), which analyzes the physicochemical evolution of proteins based on evolutionary pathways and therefore demonstrates higher accuracy in the computation of selection at non-synonymous sites than dN/dS calculations (McClellan et al., 2005; Taylor et al. 2005; PérezLosada et al. 2005). The implications of these data can help demonstrate functionally important sites including protein-protein interaction residues, etc. Residues found to be under selection in previous studies have demonstrated unique ways of inferring possible functions of unknown regions a priori, and have supported previous biochemical studies on regions of functional importance (McClellan et al., 2005). These data may then be potentially exploited in the development of biomedical procedures and pharmaceuticals, leading to improved treatment of neuromuscular and neurodegenerative diseases. In general, such physicochemical analyses allow for the (1) diagnosis of protein regions not previously understood, and (2) inference of structural and functional information of biological importance. Materials and Methods Collection of Sequence Data. Initial work began by collecting protein-coding DNA sequences from 28 taxa across arthropods and vertebrates. Sequences were collected for the α subunit gene on the NCBI (National Center for Biotechnological Information) GenBank Database. These data were then imported to the Alignment Explorer function of MEGA3 (Kumar et al. 2004). Sequence Alignment and Tree Reconstruction. Alignments were performed using ClustalW (Thompson et al. 1999) and exported in NEXUS format. After sequence alignment, appropriate models of evolution were tested and selected via PAUP* (Swofford, 2002) using a ModelTest block. The resulting scores file was then analyzed using ModelTest (Posada, 1998) and the pertinent data extracted and appended to a PAUP* block. PAUP* was used to reconstruct a maximum likelihood phylogenetic tree. TreeSAAP Analysis. The overall logic of the analyses used in this study is illustrated in Figure 1. TreeSAAP utilizes baseml (part of PAML package; Yang 1997) to reconstruct likely ancestral states for gene sequences, and then assigns weight values Figure 1. Algorithmic flow-chart describing the implementation of the MM01 model in TreeSAAP version 3: (1) nucleotide characters are optimized onto a well-corroborated phylogenetic tree to (2) infer amino acid replacement events and evolutionary pathways to define the global null hypothesis; (3) observed amino acid replacements are analyzed in the context of the several physicochemical properties to determine the number of changes per magnitude class for each amino acid index; (4) the codon compositions of the extant DNA sequences are analyzed in terms of relative frequencies of magnitude classes of evolutionary pathways for each amino acid property; (5) the fit of the expected and observed distributions are calculated using a likelihood ratio goodness of fit test, and the relative deviation of the ratio of observed to expected for each magnitude category is estimated using a z-score test to identify those properties that, on average, have been affected by positive selection; (6) a sliding-window analysis, also implementing a z-score test, is used to analyze local regions of the data set for positive selection relative to each property (critical value with Bonferroni correction illustrated as solid line and areas affected by positive selection circled); and (7) sliding-window results are correlated back to the amino acid replacements inferred from the phylogenetic tree to determine which substitutions were likely to have been affected by positive selection relative to each property. These selected changes, once identified can be correlated with aspects of protein structure and phylogeny for a more robust biological interpretation of the analytical results. to nonsynonymous codon changes and the effects of these changes in established physicochemical properties found on the AAindex (Kawashima and Kanehisa 2000). The null model to which comparisons are made is based on neutral theory (Kimura, 1991), and assumes that all or most mutations/changes will be selectively neutral or nearly so. Amino acid changing substitutions are weighted on the basis of the magnitude of changes in 31 physicochemical amino acid properties. Significant statistical deviations from the null model are interpreted as being the effect of natural selection. If the rate of certain magnitudes of change is more frequent that expected due to chance these changes are said to be the result of positive selection. When these magnitudes of change are the result of radical amino acid replacements, positive selection is said to be destabilizing and the result of molecular adaptation. A sliding window is used as a filter to identify local regions where similar amino acid replacements cluster together, and functions in the elimination of data that are exclusively the result of neutral changes (noise). Biological Annotation. Phylogenies, and primary, secondary, and tertiary structures of the proteins in question were annotated to represent the inferences made using TreeSAAP. Such a representation makes further inferences about potential interaction with the surrounding environment possible by presenting selected sites in a protein-structure context. Annotation of selected sites on a phylogeny allows visualization of a relative chronology of significant evolutionary events in the context of the pattern of speciation. Construction of sliding window graphs illustrate the spatial dynamics of selection across the primary sequence of the gene (e.g., Fig 1, panel 6). Two-dimensional secondary structure representations were annotated (not shown) to evaluate selected sites in the context of their relative location within plasma membrane. Finally, three-dimensional representations of the protein were modeled using PyMol (Delano, 2002), with significant residues corresponding to the annotated tree highlighted (Fig 2).
ABSTRACT Seventy-two full genomes corresponding to nine mammalian (67 strains) and two avian (5 strains) polyomavirus species were analyzed using maximum likelihood and Bayesian methods of phylogenetic inference. Our fully resolved and well-supported (bootstrap proportions > 90%; posterior probabilities = 1.0) trees separate the bird polyomaviruses (avian polyomavirus and goose hemorrhagic polyomavirus) from the mammalian polyomaviruses, which supports the idea of spitting the genus into two subgenera. Such a split is also consistent with the different viral life strategies of each group. Simian (simian virus 40, simian agent 12 [Sa12], and lymphotropic polyomavirus) and rodent (hamster polyomavirus, mouse polyomavirus, and murine pneumotropic polyomavirus [MPtV]) polyomaviruses did not form monophyletic groups. Using our best hypothesis of polyomavirus evolutionary relationships and established host phylogenies, we performed a cophylogenetic reconciliation analysis of codivergence. Our analyses generated six optimal cophylogenetic scenarios of coevolution, including 12 codivergence events ( P < 0.01), suggesting that Polyomaviridae coevolved with their avian and mammal hosts. As individual lineages, our analyses showed evidence of host switching in four terminal branches leading to MPtV, bovine polyomavirus, Sa12, and BK virus, suggesting a combination of vertical and horizontal transfer in the evolutionary history of the polyomaviruses.
Mark Clement合作论文数Brigham Young University in the Computer Science Department.2
Quinn O. Snell合作论文数Brigham Young University2