The evolutionary fate of proteins is driven by both folding stability and biological function, dual constraints that often conflict, creating frustration and imposing functional costs beyond stability. These costs can be captured by a "dark energy": the difference between the evolutionary energy of protein sequences and their physical folding energy. Recent advances in deep mutational scanning, protein language models, and inverse-folding models have enabled the quantification of dark energy across the protein universe. We review the computational and experimental approaches that disentangle folding and function at scale, revealing a dark energy component and providing new insights into how biological information flows from sequence to structure to function and back to sequence.
We show how to localize and quantify the functional evolutionary constraints on natural proteins. Protein folding has been one of the strongest constraints in sequence evolution. The method we propose compares the perturbations caused by local sequence variants to the energetics of the protein folding process and to the corresponding change to the apparent selection landscape of sequences over the evolutionary time scale. The difference between the physical folding free energies and the evolutionary free energies can be called a "dark energy." We analyze various protein sets and thereby show that dark energy is largely localized at functional sites, which are also often energetically frustrated from the point of view of folding. Overall, we find that about 25% of the positions of the folded globular proteins display some significant dark energy. When a function relies on a free energy that can be thermodynamically quantified, such as the binding energy to a partner, the relationship of this physical free energy with dark energy can be used to define a functional selection temperature, just as there is a selection temperature for folding. We show that selection for folding and binding functions bear similar weights in specific protein-protein interactions.
Intrinsically disordered protein regions (IDRs) mediate key steps in cellular signaling but the nature of the energy barriers crossed by IDRs upon association and the energetic features of the transition state ensemble (TSE) for binding remain elusive. Short linear motifs (SLiMs) are small functional units found within IDRs that fold upon binding to globular domains. Here, we use the LxCxE SLiM from the human papillomavirus E7 protein (LxCxEWT) binding to the retinoblastoma (Rb) protein as a minimal model system for an IDR-domain interaction. We combine extensive mutagenesis with electrostatic dissection to gain information on the binding TSE energetics and rule out ground state effects using Far-UV CD and NMR. This approach uncovers strong compensatory energetics whereby stabilizing electrostatic interactions play a key role buffering multiple weakly destabilizing non-native interactions in the binding TSE for LxCxEWT. Electrostatic buffering may enable a dynamic search for native contacts and fine tuning of the TSE energetics. A global analysis of 177 mutations reveals non-native energetics in the binding TSE of many IDRs, suggesting that electrostatic buffering may be a widespread phenomenon that dictates functional selection for charge content within IDRs. ### Competing Interest Statement The authors have declared no competing interest. Agencia Nacional de Promoción Científica y Tecnológica, 2019-02119, 2021-001027 European Union Horizon Europe MSCA Staff Exchange, 778247
Encoding of protein-coding sequences in a genome through evolution leads to characteristic proportions of codons and amino acids. Here, we present a simplified maximum entropy model that groups together codons with the same GC (guanine + cytosine) content and coding for the same amino acid and accounts for the stoichiometry of genetic elements in over 50000 genomes with seven interpretable parameters. Our model includes both the cost of a codon given a genomic GC content and the metabolic cost of the corresponding amino acid. Both costs are essential for accurate prediction of codon and amino acid abundances. The best implementation of the model includes a universal equilibrium value for the genomic GC content below 50
The tumor suppressor p53 modulates the transcription of a variety of genes, constituting a protective barrier against anomalous cellular proliferation. High-frequency “hotspot” mutations result in loss of function by the formation of amyloid-like aggregates that correlate with cancerous progression. We show that full-length p53 undergoes spontaneous homotypic condensation at submicromolar concentrations and in the absence of crowders to yield dynamic coacervates that are stoichiometrically dissolved by DNA. These coacervates fuse and evolve into hydrogel-like clusters with strong thioflavin T binding capacity, which further evolve into fibrillar species with a clearcut branching growth pattern. The amyloid-like coacervates can be rescued by the human papillomavirus master regulator E2 protein to yield large regular droplets. Furthermore, we kinetically dissected an overall condensation mechanism, which consists of a nucleation-growth process by the sequential addition of p53 tetramers, leading to discretely sized and monodisperse early condensates followed by coalescence into bead-like coacervates that slowly evolve to the fibrillar species. Our results suggest strong similarities to condensation-to-amyloid transitions observed in neurological aggregopathies. Mechanistic insights uncover novel key early and intermediate stages of condensation that can be targeted for p53 rescuing drug discovery.
We propose that spontaneous folding and molecular evolution of biopolymers are two universal aspects that must concur for life to happen. These aspects are fundamentally related to the chemical composition of biopolymers and crucially depend on the solvent in which they are embedded. We show that molecular information theory and energy landscape theory allow us to explore the limits that solvents impose on biopolymer existence. We consider 54 solvents, including water, alcohols, hydrocarbons, halogenated solvents, aromatic solvents, and low molecular weight substances made up of elements abundant in the universe, which may potentially take part in alternative biochemistries. We find that along with water, there are many solvents for which the liquid regime is compatible with biopolymer folding and evolution. We present a ranking of the solvents in terms of biopolymer compatibility. Many of these solvents have been found in molecular clouds or may be expected to occur in extrasolar planets.
Although protein sequences encode the information for folding and function, understanding their link is not an easy task. Unluckily, the prediction of how specific amino acids contribute to these features is still considerably impaired. Here, we developed a simple algorithm that finds positions in a protein sequence with potential to modulate the studied quantitative phenotypes. From a few hundred protein sequences, we perform multiple sequence alignments, obtain the per-position pairwise differences for both the sequence and the observed phenotypes, and calculate the correlation between these last two quantities. We tested our methodology with four cases: archaeal Adenylate Kinases and the organisms optimal growth temperatures, microbial rhodopsins and their maximal absorption wavelengths, mammalian myoglobins and their muscular concentration, and inhibition of HIV protease clinical isolates by two different molecules. We found from 3 to 10 positions tightly associated with those phenotypes, depending on the studied case. We showed that these correlations appear using individual positions but an improvement is achieved when the most correlated positions are jointly analyzed. Noteworthy, we performed phenotype predictions using a simple linear model that links per-position divergences and differences in the observed phenotypes. Predictions are comparable to the state-of-art methodologies which, in most of the cases, are far more complex. All of the calculations are obtained at a very low information cost since the only input needed is a multiple sequence alignment of protein sequences with their associated quantitative phenotypes. The diversity of the explored systems makes our work a valuable tool to find sequence determinants of biological activity modulation and to predict various functional features for uncharacterized members of a protein family.
The α-Proteobacteria belonging to Bradyrhizobium genus are microorganisms of extreme slow growth. Despite their extended use as inoculants in soybean production, their physiology remains poorly characterized. In this work, we produced quantitative data on four different isolates: B. diazoefficens USDA110, B. diazoefficiens USDA122, B. japonicum E109 and B. japonicum USDA6 which are representative of specific genomic profiles. Notably, we found conserved physiological traits conserved in all the studied isolates: (i) the lag and initial exponential growth phases display cell aggregation; (ii) the increase in specific nutrient concentration such as yeast extract and gluconate hinders growth; (iii) cell size does not correlate with culture age; and (iv) cell cycle presents polar growth. Meanwhile, fitness, cell size and in vitro growth widely vary across isolates correlating to ribosomal RNA operon number. In summary, this study provides novel empirical data that enriches the comprehension of the Bradyrhizobium (slow) growth dynamics and cell cycle.
TAR DNA-binding protein 43 (TDP-43) proteinopathy in brain cells is the hallmark of amyotrophic lateral sclerosis (ALS) but its cause remains elusive. Asparaginase-like-1 protein (ASRGL1) cleaves isoaspartates, which alter protein folding and susceptibility to proteolysis. ASRGL1 gene harbors a copy of the human endogenous retrovirus HML-2, whose overexpression contributes to ALS pathogenesis. Here we show that ASRGL1 expression was diminished in ALS brain samples by RNA sequencing, immunohistochemistry, and western blotting. TDP-43 and ASRGL1 colocalized in neurons but, in the absence of ASRGL1, TDP-43 aggregated in the cytoplasm. TDP-43 was found to be prone to isoaspartate formation and a substrate for ASRGL1. ASRGL1 silencing triggered accumulation of misfolded, fragmented, phosphorylated and mislocalized TDP-43 in cultured neurons and motor cortex of female mice. Overexpression of ASRGL1 restored neuronal viability. Overexpression of HML-2 led to ASRGL1 silencing. Loss of ASRGL1 leading to TDP-43 aggregation may be a critical mechanism in ALS pathophysiology.
Nearly 100 years ago, Winogradsky published a classic communication in which he described two groups of microbes, zymogenic and autochthonous. When organic matter penetrates the soil, zymogenic microbes quickly multiply and degrade it, then giving way to the slow combustion of autochthonous microbes. Although the text was originally written in French, it is often cited by English-speaking authors. We undertook a complete translation of the 1924 publication, which we provide as Supporting information. Here, we introduce the translation and describe how the zymogenic/autochthonous dichotomy shaped research questions in the study of microbial diversity and physiology. We also identify in the literature three additional and closely related dichotomies, which we propose to call exclusive copiotrophs/oligotrophs, coexisting copiotrophs/oligotrophs and fast-growing/slow-growing microbes. While Winogradsky focussed on a successional view of microbial populations over time, the current discussion is focussed on the differences in the specific growth rate of microbes as a function of the concentration of a given limiting substrate. In the future, it will be relevant to keep in mind both nutrient-focussed and time-focussed microbial dichotomies and to design experiments with both isolated laboratory cultures and multi-species communities in the spirit of Winogradsky's direct method.
Microbes are often discussed in terms of dichotomies such as copiotrophic/oligotrophic and fast/slow-growing microbes, defined using the characterisation of microbial growth in isolated cultures. The dichotomies are usually qualitative and/or study-specific, sometimes precluding clear-cut results interpretation. We can unravel microbial dichotomies as life history strategies by combining ecology theory with Monod curves, a laboratory mathematical tool of bacterial physiology that relates the specific growth rate of a microbe with the concentration of a limiting nutrient. Fitting of Monod curves provides quantities that directly correspond to key parameters in ecological theories addressing species coexistence and diversity, such as r/K selection theory, resource competition and community structure theory and the CSR triangle of life strategies. The resulting model allows us to reconcile the copiotrophic/oligotrophic and fast/slow-growing dichotomies as different subsamples of a life history strategy triangle that also includes r/K strategists. We also used the number of known carbon sources together with community structure theory to partially explain the diversity of heterotrophic microbes observed in metagenomics experiments. In sum, we propose a theoretical framework for the study of natural microbial communities that unifies several existing proposals. Its application would require the integration of metagenomics, metametabolomics, Monod curves and carbon source data.
The guanine/cytosine (GC) content of prokaryotic genomes is species-specific, taking values from 16% to 77%. This diversity of selection for GC content remains contentious. We analyse the correlations between GC content and a range of phenotypic and genotypic data in thousands of prokaryotes. GC content integrates well with these traits into r/K selection theory when phenotypic plasticity is considered. High GC-content prokaryotes are r-strategists with cheaper descendants thanks to a lower average amino acid metabolic cost, colonize unstable environments thanks to flagella and a bacillus form and are generalists in terms of resource opportunism and their defence mechanisms. Low GC content prokaryotes are K-strategists specialized for stable environments that maintain homeostasis via a high-cost outer cell membrane and endospore formation as a response to nutrient deprivation, and attain a higher nutrient-to-biomass yield. The lower proteome cost of high GC content prokaryotes is driven by the association between GC-rich codons and cheaper amino acids in the genetic code, while the correlation between GC content and genome size may be partly due to functional diversity driven by r/K selection. In all, molecular diversity in the GC content of prokaryotes may be a consequence of ecological r/K selection.
Over one hundred Mastadenovirus types infect seven orders of mammals. Virus-host coevolution may involve cospeciation, duplication, host switch and partial extinction events. We reconstruct Mastadenovirus diversification, finding that while cospeciation is dominant, the other three events are also common in Mastadenovirus evolution. Linear motifs are fast-evolving protein functional elements and key mediators of virus-host interactions, thus likely to partake in adaptive viral evolution. We study the evolution of eleven linear motifs in the Mastadenovirus E1A protein, a hub of virus-host protein–protein interactions, in the context of host diversification. The reconstruction of linear motif gain and loss events shows fast linear motif turnover, corresponding a virus-host protein–protein interaction turnover orders of magnitude faster than in model host proteomes. Evolution of E1A linear motifs is coupled, indicating functional coordination at the protein scale, yet presents motif-specific patterns suggestive of convergent evolution. We report a pervasive association between Mastadenovirus host diversification events and the evolution of E1A linear motifs. Eight of 17 host switches associate with the gain of one linear motif and the loss of four different linear motifs, while five of nine partial extinctions associate with the loss of one linear motif. The specific changes in E1A linear motifs during a host switch or a partial extinction suggest that changes in the host molecular environment lead to modulation of the interactions with the retinoblastoma protein and host transcriptional regulators. Altogether, changes in the linear motif repertoire of a viral hub protein are associated with adaptive evolution events during Mastadenovirus evolution.
Asparagines in proteins deamidate spontaneously, which changes the chemical structure of a protein and often affects its function. Current prediction algorithms for asparagine deamidation require a structure as an input or are too slow to be applied at a proteomic scale. We present NGOME-Lite, a new version of our sequence-based predictor for spontaneous asparagine deamidation that is faster by over two orders of magnitude at a similar degree of accuracy. The algorithm takes into account intrinsic sequence propensities and slowing down of deamidation by local structure. NGOME-Lite can run in a proteomic analysis mode that provides the half-time of the intact form of each protein, predicted by taking into account sequence propensities and structural protection or sequence propensities only, and a structure protection factor. The detailed analysis mode also provides graphical output for all Asn residues in the query sequence. We applied NGOME-Lite to over 257,000 sequences in 38 proteomes and found that different taxa differ in their predicted deamidation dynamics. Spontaneous protein deamidation is faster in Eukarya than in Bacteria because of a higher degree of structural protection in the latter. Predicted protein deamidation half-lifes correlate with protein turnover in human, mouse, rat, C. elegans and budding yeast but not in two plants and two bacteria. NGOME-Lite is implemented in a docker container available at https://ngome.proteinphysiologylab.org.
We study the limits imposed by transcription factor specificity on the maximum number of binding motifs that can coexist in a gene regulatory network, using the SwissRegulon Fantom5 collection of 684 human transcription factor binding sites as a model. We describe transcription factor specificity using regular expressions and find that most human transcription factor binding site motifs are separated in sequence space by one to three motif-discriminating positions. We apply theorems based on the pigeonhole principle to calculate the maximum number of transcription factors that can coexist given this degree of specificity, which is in the order of ten thousand and would fully utilize the space of DNA subsequences. Taking into account an expanded DNA alphabet with modified bases can further raise this limit by several orders of magnitude, at a lower level of sequence space usage. Our results may guide the design of transcription factors at both the molecular and system scale.
We propose an application of molecular information theory to analyze the folding of single domain proteins. We analyze results from various areas of protein science, such as sequence-based potentials, reduced amino acid alphabets, backbone configurational entropy, secondary structure content, residue burial layers, and mutational studies of protein stability changes. We found that the average information contained in the sequences of evolved proteins is very close to the average information needed to specify a fold ∼2.2 ± 0.3 bits/(site·operation). The effective alphabet size in evolved proteins equals the effective number of conformations of a residue in the compact unfolded state at around 5. We calculated an energy-to-information conversion efficiency upon folding of around 50%, lower than the theoretical limit of 70%, but much higher than human-built macroscopic machines. We propose a simple mapping between molecular information theory and energy landscape theory and explore the connections between sequence evolution, configurational entropy, and the energetics of protein folding.
Although protein sequences encode the information for folding and function, understanding their link is not an easy task. Unluckily, the prediction of how specific amino acids contribute to these features is still considerably impaired. Here, we developed PhISCO, Phenotype Inference from Sequence COmparisons, a simple algorithm that finds positions associated with any quantitative phenotype and predicts their values. From a few hundred sequences from four different protein families, we performed multiple sequence alignments and calculated per-position pairwise differences for both the sequence and the observed phenotypes. We found that from 3 to 10 positions, depending on the studied case, were enough to identify positions associated with the phenotypes and perform quantitative predictions of them. Here we show that these strong correlations can be found using individual positions while an improvement is achieved when the most correlated positions are jointly analyzed. Noteworthy, we performed phenotype predictions using a simple linear model that links per-position divergences and differences in observed phenotypes. We also show that although extremely simple, predictions are comparable to the state-of-art methodologies which, in most of the cases, are far more complex. All of the calculations are obtained at a very low information cost since the only input needed is a multiple sequence alignment of protein sequences with their associated quantitative phenotype. The diversity of the explored systems makes PhISCO a valuable tool to find sequence determinants of biological activity modulation and to predict various functional features for uncharacterized members of a protein family.
Many disordered proteins conserve essential functions in the face of extensive sequence variation, making it challenging to identify the mechanisms responsible for functional selection. Here we identify the molecular mechanism of functional selection for the disordered adenovirus early gene 1A (E1A) protein. E1A competes with host factors to bind the retinoblastoma (Rb) protein, subverting cell cycle regulation. We show that two binding motifs tethered by a hypervariable disordered linker drive picomolar affinity Rb binding and host factor displacement. Compensatory changes in amino acid sequence composition and sequence length lead to conservation of optimal tethering across a large family of E1A linkers. We refer to this compensatory mechanism as conformational buffering. We also detect coevolution of the motifs and linker, which can preserve or eliminate the tethering mechanism. Conformational buffering and motif-linker coevolution explain robust functional encoding within hypervariable disordered linkers and could underlie functional selection of many disordered protein regions.