We recently reanalyzed 20 combinatorial mutagenesis datasets using a novel reference-free analysis (RFA) method and showed that high-order epistasis contributes negligibly to protein sequence-function relationships in every case. Dupic, Phillips, and Desai (DPD) commented on a preprint of our work. In our published paper, we addressed all the major issues they raised, but we respond directly to them here. 1) DPD's claim that RFA is equivalent to estimating reference-based analysis (RBA) models by regression neglects fundamental differences in how the two formalisms dissect the causal relationship between sequence and function. It also misinterprets the observation that using regression to estimate any truncated model of genetic architecture will always yield the same predicted phenotypes and variance partition; the resulting estimates correspond to those of the RFA formalism but are inaccurate representations of the true RBA model. 2) DPD's claim that high-order epistasis is widespread and significant while somehow explaining little phenotypic variance is an artifact of two strong biases in the use of regression to estimate RBA models: this procedure underestimates the phenotypic variance explained by RBA epistatic terms while at the same time inflating the magnitude of individual terms. 3) DPD erroneously claim that RFA is "exactly equivalent" to Fourier analysis (FA) and background-averaged analysis (BA). This error arises because DPD used an incorrect mathematical definition of RFA and were misled by a simple numerical relationship among the models that only holds only for the simplest kinds of datasets. 4) DPD argue that using a nonlinear transformation to account for global nonlinearities in sequence-function relationships is often unnecessary and may artifactually absorb specific epistatic interactions. We show that nonspecific epistasis caused by a limited dynamic range affects datasets of all types, even when the phenotype is represented on a free-energy scale. Moreover, using a nonlinear transformation in a joint fitting procedure does not underestimate specific epistasis under realistic conditions, even if the data are not affected by nonspecific epistasis. The conclusions of our work therefore hold: the genetic architecture of all 20 protein datasets we analyzed can be efficiently and accurately described in an RFA framework by first-order amino acid effects and pairwise interactions with a simple model of global nonlinearity. We are grateful for DPD's commentary, which helped us improve our paper.
A protein's genetic architecture – the set of causal rules by which its sequence produces its functions – also determines its possible evolutionary trajectories. Prior research has proposed that genetic architecture of proteins is very complex, with pervasive epistatic interactions that constrain evolution and make function difficult to predict from sequence. Most of this work has analyzed only the direct paths between two proteins of interest – excluding the vast majority of possible genotypes and evolutionary trajectories – and has considered only a single protein function, leaving unaddressed the genetic architecture of functional specificity and its impact on the evolution of new functions. Here we develop a new method based on ordinal logistic regression to directly characterize the global genetic determinants of multiple protein functions from 20-state combinatorial deep mutational scanning (DMS) experiments. We use it to dissect the genetic architecture and evolution of a transcription factor's specificity for DNA, using data from a combinatorial DMS of an ancient steroid hormone receptor's capacity to activate transcription from two biologically relevant DNA elements. We show that the genetic architecture of DNA recognition consists of a dense set of main and pairwise effects that involve virtually every possible amino acid state in the protein-DNA interface, but higher-order epistasis plays only a tiny role. Pairwise interactions enlarge the set of functional sequences and are the primary determinants of specificity for different DNA elements. They also massively expand the number of opportunities for single-residue mutations to switch specificity from one DNA target to another. By bringing variants with different functions close together in sequence space, pairwise epistasis therefore facilitates rather than constrains the evolution of new functions.
A protein’s genetic architecture – the set of causal rules by which its sequence produces its functions – also determines its possible evolutionary trajectories. Prior research has proposed that the genetic architecture of proteins is very complex, with pervasive epistatic interactions that constrain evolution and make function difficult to predict from sequence. Most of this work has analyzed only the direct paths between two proteins of interest – excluding the vast majority of possible genotypes and evolutionary trajectories – and has considered only a single protein function, leaving unaddressed the genetic architecture of functional specificity and its impact on the evolution of new functions. Here, we develop a new method based on ordinal logistic regression to directly characterize the global genetic determinants of multiple protein functions from 20-state combinatorial deep mutational scanning (DMS) experiments. We use it to dissect the genetic architecture and evolution of a transcription factor’s specificity for DNA, using data from a combinatorial DMS of an ancient steroid hormone receptor’s capacity to activate transcription from two biologically relevant DNA elements. We show that the genetic architecture of DNA recognition consists of a dense set of main and pairwise effects that involve virtually every possible amino acid state in the protein-DNA interface, but higher-order epistasis plays only a tiny role. Pairwise interactions enlarge the set of functional sequences and are the primary determinants of specificity for different DNA elements. They also massively expand the number of opportunities for single-residue mutations to switch specificity from one DNA target to another. By bringing variants with different functions close together in sequence space, pairwise epistasis therefore facilitates rather than constrains the evolution of new functions.
Many enzymes assemble into homomeric protein complexes comprising multiple copies of one protein. Because structural form is usually assumed to follow function in biochemistry, these assemblies are thought to evolve because they provide some functional advantage. In many cases, however, no specific advantage is known and, in some cases, quaternary structure varies among orthologs. This has led to the proposition that self-assembly may instead vary neutrally within protein families. The extent of such variation has been difficult to ascertain because quaternary structure has until recently been difficult to measure on large scales. Here, we employ mass photometry, phylogenetics, and structural biology to interrogate the evolution of homo-oligomeric assembly across the entire phylogeny of prokaryotic citrate synthases - an enzyme with a highly conserved function. We discover a menagerie of different assembly types that come and go over the course of evolution, including cases of parallel evolution and reversions from complex to simple assemblies. Functional experiments in vitro and in vivo indicate that evolutionary transitions between different assemblies do not strongly influence enzyme catalysis. Our work suggests that enzymes can wander relatively freely through a large space of possible assembly states and demonstrates the power of characterizing structure-function relationships across entire phylogenies.
The relationship between genetic code robustness and protein evolvability is unknown. A new study in PLOS Biology using in silico rewiring of genetic codes and functional protein data identified a positive correlation between code robustness and protein evolvability that is protein-specific.
How complex are the rules by which a protein's sequence determines its function? High-order epistatic interactions among residues are thought to be pervasive, suggesting an idiosyncratic and unpredictable sequence-function relationship. But many prior studies may have overestimated epistasis, because they analyzed sequence-function relationships relative to a single reference sequence-which causes measurement noise and local idiosyncrasies to snowball into high-order epistasis-or they did not fully account for global nonlinearities. Here we present a reference-free method that jointly infers specific epistatic interactions and global nonlinearity using a bird's-eye view of sequence space. This technique yields the simplest explanation of sequence-function relationships and is more robust than existing methods to measurement noise, missing data, and model misspecification. We reanalyze 20 experimental datasets and find that context-independent amino acid effects and pairwise interactions, along with a simple nonlinearity to account for limited dynamic range, explain a median of 96% of phenotypic variance and over 92% in every case. Only a tiny fraction of genotypes are strongly affected by higher-order epistasis. Sequence-function relationships are also sparse: a miniscule fraction of amino acids and interactions account for 90% of phenotypic variance. Sequence-function causality across these datasets is therefore simple, opening the way for tractable approaches to characterize proteins' genetic architecture. Understanding protein sequence-function relationships is complicated by high order epistatic interactions among residues, although the extent of these interactions remains uncertain. Here, the authors present a reference-free method which suggests that sequence-function relationships are relatively simple, with little influence from high order epistatic interactions.
Epistatic interactions can make the outcomes of evolution unpredictable, but no comprehensive data are available on the extent and temporal dynamics of changes in the effects of mutations as protein sequences evolve. Here, we use phylogenetic deep mutational scanning to measure the functional effect of every possible amino acid mutation in a series of ancestral and extant steroid receptor DNA binding domains. Across 700 million years of evolution, epistatic interactions caused the effects of most mutations to become decorrelated from their initial effects and their windows of evolutionary accessibility to open and close transiently. Most effects changed gradually and without bias at rates that were largely constant across time, indicating a neutral process caused by many weak epistatic interactions. Our findings show that protein sequences drift inexorably into contingency and unpredictability, but that the process is statistically predictable, given sufficient phylogenetic and experimental data.
Heritable variation in a gene’s expression arises from mutations impacting cis - and trans -acting components of its regulatory network. Here, we investigate how trans -regulatory mutations are distributed within the genome and within a gene regulatory network by identifying and characterizing 69 mutations with trans -regulatory effects on expression of the same focal gene in Saccharomyces cerevisiae . Relative to 1766 mutations without effects on expression of this focal gene, we found that these trans -regulatory mutations were enriched in coding sequences of transcription factors previously predicted to regulate expression of the focal gene. However, over 90% of the trans -regulatory mutations identified mapped to other types of genes involved in diverse biological processes including chromatin state, metabolism, and signal transduction. These data show how genetic changes in diverse types of genes can impact a gene’s expression in trans , revealing properties of trans -regulatory mutations that provide the raw material for trans -regulatory variation segregating within natural populations.
The roles of chance, contingency, and necessity in evolution are unresolved because they have never been assessed in a single system or on timescales relevant to historical evolution. We combined ancestral protein reconstruction and a new continuous evolution technology to mutate and select proteins in the B-cell lymphoma-2 (BCL-2) family to acquire protein–protein interaction specificities that occurred during animal evolution. By replicating evolutionary trajectories from multiple ancestral proteins, we found that contingency generated over long historical timescales steadily erased necessity and overwhelmed chance as the primary cause of acquired sequence variation; trajectories launched from phylogenetically distant proteins yielded virtually no common mutations, even under strong and identical selection pressures. Chance arose because many sets of mutations could alter specificity at any timepoint; contingency arose because historical substitutions changed these sets. Our results suggest that patterns of variation in BCL-2 sequences – and likely other proteins, too – are idiosyncratic products of a particular and unpredictable course of historical events.
Most proteins assemble into multisubunit complexes1. The persistence of these complexes across evolutionary time is usually explained as the result of natural selection for functional properties that depend on multimerization, such as intersubunit allostery or the capacity to do mechanical work2. In many complexes, however, multimerization does not enable any known function3. An alternative explanation is that multimers could become entrenched if substitutions accumulate that are neutral in multimers but deleterious in monomers; purifying selection would then prevent reversion to the unassembled form, even if assembly per se does not enhance biological function3-7. Here we show that a hydrophobic mutational ratchet systematically entrenches molecular complexes. By applying ancestral protein reconstruction and biochemical assays to the evolution of steroid hormone receptors, we show that an ancient hydrophobic interface, conserved for hundreds of millions of years, is entrenched because exposure of this interface to solvent reduces protein stability and causes aggregation, even though the interface makes no detectable contribution to function. Using structural bioinformatics, we show that a universal mutational propensity drives sites that are buried in multimeric interfaces to accumulate hydrophobic substitutions to levels that are not tolerated in monomers. In a database of hundreds of families of multimers, most show signatures of long-term hydrophobic entrenchment. It is therefore likely that many protein complexes persist because a simple ratchet-like mechanism entrenches them across evolutionary time, even when they are functionally gratuitous.
The extent to which evolutionary outcomes reflect the unpredictable influences of chance and contingency is a central but unanswered question in evolutionary biology . A precise characterization requires evolutionary trajectories to be repeated multiple times under identical environmental conditions from multiple starting points across history, a scenario that rarely, if ever, occurs in nature. Here we combine continuous experimental evolution with ancestral protein reconstruction and manipulative genetic experiments to identify the causes and consequences of chance and contingency in the genetic outcomes of molecular evolution. By repeatedly evolving ancestral proteins in the B-cell lymphoma-2 (BCL-2) family of apoptosis regulators to acquire the same protein-protein interaction specificities that evolved during history, we found that contingency and chance interact to make sequence evolution increasingly unpredictable over phylogenetic timescales. Although replicates from the same starting genotype sometimes share mutations – indicating partial predictability – there are multiple alternative sets of changes that can alter specificity, and chance decides which of these paths is taken. Contingency has a stronger effect: when trajectories are initiated from different starting points, outcomes are even more divergent, because substitutions that occurred during phylogenetic history repeatedly changed the potential of other mutations to confer new binding specificities. The impact of contingency increased steadily with phylogenetic distance and magnified the effects of chance, resulting in a >3-fold increase in genetic variance among evolutionary trajectories initiated from different starting points across the timescale of metazoan evolution. Our findings show how a particular cascade of chance evolutionary steps throughout history makes the outcomes of molecular evolution increasingly idiosyncratic and unpredictable, even under strong selection.
Heritable variation in gene expression is common within species. Much of this variation is due to genetic differences outside of the gene with altered expression and is trans-acting. This trans-regulatory variation is often polygenic, with individual variants typically having small effects, making the genetic architecture and evolution of trans-regulatory variation challenging to study. Consequently, key questions about trans-regulatory variation remain, including the variability of trans-regulatory variation within a species, how selection affects trans-regulatory variation, and how trans-regulatory variants are distributed throughout the genome and within a species. To address these questions, we isolated and measured trans-regulatory differences affecting TDH3 promoter activity among 56 strains of Saccharomyces cerevisiae, finding that trans-regulatory backgrounds varied approximately twofold in their effects on TDH3 promoter activity. Comparing this variation to neutral models of trans-regulatory evolution based on empirical measures of mutational effects revealed that despite this variability in the effects of trans-regulatory backgrounds, stabilizing selection has constrained trans-regulatory differences within this species. Using a powerful quantitative trait locus mapping method, we identified ∼100 trans-acting expression quantitative trait locus in each of three crosses to a common reference strain, indicating that regulatory variation is more polygenic than previous studies have suggested. Loci altering expression were located throughout the genome, and many loci were strain specific. This distribution and prevalence of alleles is consistent with recent theories about the genetic architecture of complex traits. In all mapping experiments, the nonreference strain alleles increased and decreased TDH3 promoter activity with similar frequencies, suggesting that stabilizing selection maintained many trans-acting variants with opposing effects. This variation may provide the raw material for compensatory evolution and larger scale regulatory rewiring observed in developmental systems drift among species.
Heritable variation in gene expression is common within species. Much of this variation is due to genetic changes at loci other than the affected gene and is thus trans -acting. This trans -regulatory variation is often polygenic, with individual variants typically having small effects, making the genetic architecture of trans -regulatory variation challenging to study. Consequently, key questions about trans -regulatory variation remain, including how selection affects this variation and how trans -regulatory variants are distributed throughout the genome and within species. Here, we show that trans -regulatory variation affecting TDH3 promoter activity is common among strains of Saccharomyces cerevisiae . Comparing this variation to neutral models of trans -regulatory evolution based on empirical measures of mutational effects revealed that stabilizing selection has constrained this variation. Using a powerful quantitative trait locus (QTL) mapping method, we identified ∼100 loci altering expression between a reference strain and each of three genetically distinct strains. In all three cases, the non-reference strain alleles increased and decreased TDH3 promoter activity with similar frequencies, suggesting that stabilizing selection maintained many trans -acting variants with opposing effects. Loci altering expression were located throughout the genome, with many loci being strain specific and others being shared among multiple strains. These findings are consistent with theory showing stabilizing selection for quantitative traits can maintain many alleles with opposing effects, and the wide-spread distribution of QTL throughout the genome is consistent with the omnigenic model of complex trait variation. Furthermore, the prevalence of alleles with opposing effects might provide raw material for compensatory evolution and developmental systems drift.Significance statement Gene expression varies among individuals in a population due to genetic differences in regulatory components. To determine how this variation is distributed within genomes and species, we used a powerful genetic mapping approach to examine multiple strains of Saccharomyces cerevisiae . Despite evidence of stabilizing selection maintaining gene expression levels among strains, we find hundreds of loci that affect expression of a single gene. These loci vary among strains and include similar frequencies of alleles that increase and decrease expression. As a result, each strain contains a unique set of compensatory alleles that lead to similar levels of gene expression among strains. This regulatory variation might form the basis for large scale regulatory rewiring observed between distantly related species.
Gene expression noise is an evolvable property of biological systems that describes differences in expression among genetically identical cells in the same environment. Prior work has shown that expression noise is heritable and can be shaped by selection, but the impact of variation in expression noise on organismal fitness has proven difficult to measure. Here, we quantify the fitness effects of altering expression noise for the TDH3 gene in Saccharomyces cerevisiae. We show that increases in expression noise can be deleterious or beneficial depending on the difference between the average expression level of a genotype and the expression level maximizing fitness. We also show that a simple model relating single-cell expression levels to population growth produces patterns consistent with our empirical data. We use this model to explore a broad range of average expression levels and expression noise, providing additional insight into the fitness effects of variation in expression noise.
Phenotypic plasticity is an evolvable property of biological systems that can arise from environment-specific regulation of gene expression. To better understand the evolutionary and molecular mechanisms that give rise to plasticity in gene expression, we quantified the effects of 235 single-nucleotide mutations in the Saccharomyces cerevisiae TDH3 promoter (P-TDH3) on the activity of this promoter in media containing glucose, galactose, or glycerol as a carbon source. We found that the distributions of mutational effects differed among environments because many mutations altered the plastic response exhibited by the wild-type allele. Comparing the effects of these mutations with the effects of 30 P-TDH3 polymorphisms on expression plasticity in the same environments provided evidence of natural selection acting to prevent the plastic response in P-TDH3 activity between glucose and galactose from becoming larger. The largest changes in expression plasticity were observed between fermentable (glucose or galactose) and nonfermentable (glycerol) carbon sources and were caused by mutations located in the RAP1 and GCR1 transcription factor binding sites. Mutations altered expression plasticity most frequently between the two fermentable environments, with mutations causing significant changes in plasticity between glucose and galactose distributed throughout the promoter, suggesting they might affect chromatin structure. Taken together, these results provide insight into the molecular mechanisms underlying gene-by-environment interactions affecting gene expression as well as the evolutionary dynamics affecting natural variation in plasticity of gene expression.
Heritable changes in gene expression are important contributors to phenotypic differences within and between species and are caused by mutations in cis-regulatory elements and trans-regulatory factors. Although previous work has suggested that cis-regulatory differences preferentially accumulate with time, technical restrictions to closely related species and limited comparisons have made this observation difficult to test. To address this problem, we used allele-specific RNA-seq data from Saccharomyces species and hybrids to expand both the evolutionary timescale and number of species in which the evolution of regulatory divergence has been investigated. We find that as sequence divergence increases, cis-regulatory differences do indeed become the dominant type of regulatory difference between species, ultimately becoming a better predictor of expression divergence than trans-regulatory divergence. When both cis- and trans-regulatory differences accumulate for the same gene, they more often have effects in opposite directions than in the same direction, indicating widespread compensatory changes underlying the evolution of gene expression. The frequency of compensatory changes within and between species and the magnitude of effect for the underlying cis- and trans-regulatory differences suggests that compensatory changes accumulate primarily due to selection against divergence in gene expression as a result of weak stabilizing selection on gene expression levels. These results show that cis-regulatory differences and compensatory changes in regulation play increasingly important roles in the evolution of gene expression as time increases.
The budding yeast Saccharomyces cerevisiae is the best studied eukaryote in molecular and cell biology, but its utility for understanding the genetic basis of phenotypic variation in natural populations is limited by inefficient association mapping due to strong and complex population structure. To overcome this challenge, we generated genome sequences for 85 strains and performed a comprehensive population genomic survey of a total of 190 diverse strains. We identified considerable variation in population structure among chromosomes and identified 181 genes that are absent from the reference genome. Many of these nonreference genes are expressed and we functionally confirmed that two of these genes confer increased resistance to antifungals. Next, we simultaneously measured the growth rates of over 4,500 laboratory strains, each of which lacks a nonessential gene, and 81 natural strains across multiple environments using unique DNA barcode present in each strain. By combining the genome-wide reverse genetic information gained from the gene deletion strains with a genome-wide association analysis from the natural strains, we identified genomic regions associated with fitness variation in natural populations. To experimentally validate a subset of these associations, we used reciprocal hemizygosity tests, finding that while the combined forward and reverse genetic approaches can identify a single causal gene, the phenotypic consequences of natural genetic variation often follow a complicated pattern. The resources and approach provided outline an efficient and reliable route to association mapping in yeast and significantly enhance its value as a model for understanding the genetic mechanisms underlying phenotypic variation and evolution in natural populations.
Heritable differences in gene expression are caused by mutations in DNA sequences encoding cis-regulatory elements and trans-regulatory factors. These two classes of regulatory change differ in their relative contributions to expression differences in natural populations because of the combined effects of mutation and natural selection. Here, we investigate how new mutations create the regulatory variation upon which natural selection acts by quantifying the frequencies and effects of hundreds of new cis- and trans-acting mutations altering activity of the TDH3 promoter in the yeast Saccharomyces cerevisiae in the absence of natural selection. We find that cis-regulatory mutations have larger effects on expression than trans-regulatory mutations and that while trans-regulatory mutations are more common overall, cis- and trans-regulatory changes in expression are equally abundant when only the largest changes in expression are considered. In addition, we find that cis-regulatory mutations are skewed toward decreased expression while trans-regulatory mutations are skewed toward increased expression. We also measure the effects of cis- and trans-regulatory mutations on the variability in gene expression among genetically identical cells, a property of gene expression known as expression noise, finding that trans-regulatory mutations are much more likely to decrease expression noise than cis-regulatory mutations. Because new mutations are the raw material upon which natural selection acts, these differences in the frequencies and effects of cis- and trans-regulatory mutations should be considered in models of regulatory evolution.
The budding yeast Saccharomyces cerevisiae is the best studied eukaryote in molecular and cell biology, but its utility for understanding the genetic basis of natural phenotypic variation is limited by the inefficiency of association mapping owing to strong and complex population structure. To facilitate association mapping, we analyzed 190 high-quality genomes of diverse strains, including 85 newly sequenced ones, to uncover yeast’s population structure that varies substantially among genomic regions. We identified 181 yeast genes that are absent from the reference genome and demonstrated their expression and role in important functions such as drug resistance. We then simultaneously measured the growth rates of over 4500 lab strains each deficient of a nonessential gene and 81 natural strains across multiple environments using unique DNA barcode present in each strain. We combined the genome-wide reverse genetic information with genome-wide association analysis to determine potential genomic regions of importance to environmental adaptations, and for a subset experimentally validated their role by reciprocal hemizygosity tests. The resources provided permit efficient and reliable association mapping in yeast and significantly enhances its value as a model for understanding the genetic mechanisms of phenotypic polymorphism and evolution.
Quantifying activity of cis-regulatory sequences controlling gene expression shows that selection on expression noise has a greater impact on sequence variation than selection on mean expression level. Patricia Wittkopp and colleagues measure and compare the effects of 236 mutations in a cis-regulatory region from Saccharomyces cerevisiae to the effects of natural polymorphisms observed in the same region among 85 isolates of this species. They find that selection on variability in expression among genetically identical cells appears to have had a greater effect on sequence variation than selection on mean expression level, at their promoter. This may not be because variation in expression noise affects fitness more than variation in mean expression level; rather, it may be due to differences in the distributions of mutational effects for these two phenotypes. Genetic variation segregating within a species reflects the combined activities of mutation, selection, and genetic drift. In the absence of selection, polymorphisms are expected to be a random subset of new mutations; thus, comparing the effects of polymorphisms and new mutations provides a test for selection1,2,3,4. When evidence of selection exists, such comparisons can identify properties of mutations that are most likely to persist in natural populations2. Here we investigate how mutation and selection have shaped variation in a cis-regulatory sequence controlling gene expression by empirically determining the effects of polymorphisms segregating in the TDH3 promoter among 85 strains of Saccharomyces cerevisiae and comparing their effects to a distribution of mutational effects defined by 236 point mutations in the same promoter. Surprisingly, we find that selection on expression noise (that is, variability in expression among genetically identical cells5) appears to have had a greater impact on sequence variation in the TDH3 promoter than selection on mean expression level. This is not necessarily because variation in expression noise impacts fitness more than variation in mean expression level, but rather because of differences in the distributions of mutational effects for these two phenotypes. This study shows how systematically examining the effects of new mutations can enrich our understanding of evolutionary mechanisms. It also provides rare empirical evidence of selection acting on expression noise.