Genomes contain orders of magnitude more open reading frames (ORFs) than known protein coding genes, and recent work suggests there may be unannotated proteins present in even the best studied organisms. To address this gap, we used a high throughput reverse genetic toolkit to construct precise C-terminal fusions of a reporter (and control) to >120,000 ORFs in model E. coli . We found hundreds of unannotated significant hits, and individually detected >50 novel polypeptides by western blot, including ORFs within tRNA loci. Many ORFs overlap annotated genes in the sense orientation, and we found these are likely chimeric polypeptides produced by ribosomal frameshifting. Using degron based knockdowns, we identified unannotated proteins that have putative fitness effects, and we found a novel small protein that displays phenotypes consistent with a role in the mRNA degradosome. The observation of a range of unannotated translation products should lead to better annotation and understanding of the bacterial domain of life and motivates the continued exploration of genomes broadly. ### Competing Interest Statement The authors have declared no competing interest.
The nature of standing genetic variation remains a central debate in population genetics, with differing perspectives on whether common variants are almost always neutral as suggested by neutral and nearly neutral theories or whether they can commonly have large functional and fitness effects as proposed by the balance theory. We address this question by mapping the fitness effects of over 9,000 natural variants in the Ras/PKA and TOR/Sch9 pathways-key regulators of cell proliferation in eukaryotes-across four conditions in Saccharomyces cerevisiae . While most variants are neutral in our assay, ~3,500 exhibited significant fitness effects. These non-neutral variants tend to be missense and to affect conserved, more densely packed, and less solvent-exposed protein regions. While some of these non-neutral variants are younger and rarer, and more often found in heterozygous states-consistent with purifying selection-a substantial fraction is present at high frequencies in the population, which is expected under balancing selection. Indeed, we find that variants with a positive fitness effect in our laboratory measurement show strong signs of local adaptation as they tend to be found specifically in domesticated strains isolated from human-made environments. Our findings support the view that while many common variants might be effectively neutral, a significant proportion have locally adaptive functional consequences and are driven into a subset of the population by local positive selection. This study highlights the potential to combine high-throughput precision genome editing with fitness measurements to explore natural genetic variation on a pathway-wide scale, thereby bridging the gap between population genetics and functional genomics to understand the nature of evolutionary forces in the wild.
Elucidating the complex relationships between genotypes, phenotypes, and fitness remains one of the fundamental challenges in evolutionary biology. Part of the difficulty arises from the enormous number of possible genotypes and the lack of understanding of the underlying phenotypic differences driving adaptation. Here, we present a computational method that takes advantage of modern high-throughput fitness measurements to learn a map from high-dimensional fitness profiles to a low-dimensional latent space in a geometry-informed manner. We demonstrate that our approach using a Riemannian Hamiltonian Variational Autoencoder (RHVAE) outperforms traditional linear dimensionality reduction techniques by capturing the nonlinear structure of the phenotype-fitness map. When applied to simulated adaptive dynamics, we show that the learned latent space retains information about the underlying adaptive phenotypic space and accurately reconstructs complex fitness landscapes. We then apply this method to a dataset of high-throughput fitness measurements of E. coli under different antibiotic pressures and demonstrate superior predictive power for out-of-sample data compared to linear approaches. Our work provides a data-driven implementation of Fisher’s geometric model of adaptation, transforming it from a theoretical framework into an empirically grounded approach for understanding evolutionary dynamics using modern deep learning methods. ### Competing Interest Statement The authors have declared no competing interest. NIH/NIGMS, R35GM11816506 (MIRA grant) NSF, DMS-2235451, PHY-1748958 Simons Foundation, MPTMPS-00005320, 597491-RWC Chan Zuckerberg Initiative (United States), DAF2023-329587 Gordon and Betty Moore Foundation, 2919.02
The tracking of lineage frequencies via DNA barcode sequencing enables the quantification of microbial fitness. However, experimental noise coming from biotic and abiotic sources complicates the computation of a reliable inference. We present a Bayesian pipeline to infer relative microbial fitness from high-throughput lineage tracking assays. Our model accounts for multiple sources of noise and propagates uncertainties throughout all parameters in a systematic way. Furthermore, using modern variational inference methods based on automatic differentiation, we are able to scale the inference to a large number of unique barcodes. We extend this core model to analyze multi-environment assays, replicate experiments, and barcodes linked to genotypes. On simulations, our method recovers known parameters within posterior credible intervals. This work provides a generalizable Bayesian framework to analyze lineage tracking experiments. The accompanying open-source software library enables the adoption of principled statistical methods in experimental evolution.
The rapid turnover of dimethylsulfoniopropionate (DMSP), likely the most relevant dissolved organic sulfur compound in the surface ocean, makes it pivotal to understand the cycling of organic sulfur. Dimethylsulfoniopropionate is mainly synthesized by phytoplankton, and it can be utilized as carbon and sulfur sources by marine bacteria or cleaved by bacteria or algae to produce the volatile compound dimethylsulfide (DMS), involved in the formation of sulfate aerosols. The fluxes between the consumption (i.e., demethylation) and cleavage pathways are thought to depend on community interactions and their sulfur demand. However, a quantitative assessment of the sulfur partitioning between each of these pathways is still missing. Here, we report for the first time the sulfur isotope fractionations by enzymes involved in DMSP degradation with different catalytic mechanisms, expressed heterologously in Escherichia coli . We show that the residual DMSP from the demethylation pathway is 2.7‰ enriched in δ 34 S relative to the initial DMSP, and that the fractionation factor ( 34 ε ) of the cleavage pathways varies between −1 and −9‰. The incorporation of these fractionation factors into mass balance calculations constrains the biological fates of DMSP in seawater, supports the notion that demethylation dominates over cleavage in marine environments, and could be used as a proxy for the dominant pathways of degradation of DMSP by marine microbial communities.
The study of transcription remains one of the centerpieces of modern biology with implications in settings from development to metabolism to evolution to disease. Precision measurements using a host of different techniques including fluorescence and sequencing readouts have raised the bar for what it means to quantitatively understand transcriptional regulation. In particular our understanding of the simplest genetic circuit is sufficiently refined both experimentally and theoretically that it has become possible to carefully discriminate between different conceptual pictures of how this regulatory system works. This regulatory motif, originally posited by Jacob and Monod in the 1960s, consists of a single transcriptional repressor binding to a promoter site and inhibiting transcription. In this paper, we show how seven distinct models of this so-called simple-repression motif, based both on thermodynamic and kinetic thinking, can be used to derive the predicted levels of gene expression and shed light on the often surprising past success of the thermodynamic models. These different models are then invoked to confront a variety of different data on mean, variance and full gene expression distributions, illustrating the extent to which such models can and cannot be distinguished, and suggesting a two-state model with a distribution of burst sizes as the most potent of the seven for describing the simple-repression motif.
Given the stochastic nature of gene expression, genetically identical cells exposed to the same environmental inputs will produce different outputs. This heterogeneity has been hypothesized to have consequences for how cells are able to survive in changing environments. Recent work has explored the use of information theory as a framework to understand the accuracy with which cells can ascertain the state of their surroundings. Yet the predictive power of these approaches is limited and has not been rigorously tested using precision measurements. To that end, we generate a minimal model for a simple genetic circuit in which all parameter values for the model come from independently published data sets. We then predict the information processing capacity of the genetic circuit for a suite of biophysical parameters such as protein copy number and protein-DNA affinity. We compare these parameter-free predictions with an experimental determination of protein expression distributions and the resulting information processing capacity of E. coli cells. We find that our minimal model captures the scaling of the cell-to-cell variability in the data and the inferred information processing capacity of our simple genetic circuit up to a systematic deviation.
Every organism has intricate regulatory networks that enable them to sense, move, and interact with complex environments. In E. coli, transcription factors (TFs) bind to short sequences upstream of genes, called operators, to repress or activate gene expression. Our lab has previously derived and validated a biophysical model of transcription regulation -the repression of a gene by LacI - demonstrating that repressorcopy number, the energy of TF:operator binding, and genome size all play crucial roles in determining the quantitative features of gene expression. But LacI is a small fish in the cellular pond; we must develop experimental methods that enable us to map how any TF:operator pair imparts predictable gene expression, and do so in a high-throughput manner without sacrificing quantitation. We report a high-throughput method to randomly mutagenize large libraries of operator sequences and quantify their resulting gene expression. Briefly, mutagenized binding sites for a TF are “mapped” to a random DNA barcode using next-generation sequencing. The native gene regulated by that TF is then inserted between the mutated operators and barcodes, and these libraries are genomically-integrated into E. coli. The relative abundance of cells carrying each operator mutant can be determined via DNA-sequencing of the barcodes, while their gene expression can be measured by quantitative RNA-sequencing of barcodes using a mixture of Unique Molecular Identifiers and other methods to minimize bias. These data will inform mathematical models for the de novo prediction of gene expression from regulatory sequences.
The study of transcription remains one of the centerpieces of modern biology with implications in settings from development to metabolism to evolution to disease. Precision measurements using a host of different techniques including fluorescence and sequencing readouts have raised the bar for what it means to quantitatively understand transcriptional regulation. In particular our understanding of the simplest genetic circuit is sufficiently refined both experimentally and theoretically that it has become possible to carefully discriminate between different conceptual pictures of how this regulatory system works. This regulatory motif, originally posited by Jacob and Monod in the 1960s, consists of a single transcriptional repressor binding to a promoter site and inhibiting transcription. In this paper, we show how seven distinct models of this so-called simple-repression motif, based both on equilibrium and kinetic thinking, can be used to derive the predicted levels of gene expression and shed light on the often surprising past success of the equilbrium models. These different models are then invoked to confront a variety of different data on mean, variance and full gene expression distributions, illustrating the extent to which such models can and cannot be distinguished, and suggesting a two-state model with a distribution of burst sizes as the most potent of the seven for describing the simple-repression motif.
Mutation is a critical mechanism by which evolution explores the functional landscape of proteins. Despite our ability to experimentally inflict mutations at will, it remains difficult to link sequence-level perturbations to systems-level responses. Here, we present a framework centered on measuring changes in the free energy of the system to link individual mutations in an allosteric transcriptional repressor to the parameters which govern its response. We find the energetic effects of the mutations can be categorized into several classes which have characteristic curves as a function of the inducer concentration. We experimentally test these diagnostic predictions using the well-characterized LacI repressor of Escherichia coli , probing several mutations in the DNA binding and inducer binding domains. We find that the change in gene expression due to a point mutation can be captured by modifying only a subset of the model parameters that describe the respective domain of the wild-type protein. These parameters appear to be insulated, with mutations in the DNA binding domain altering only the DNA affinity and those in the inducer binding domain altering only the allosteric parameters. Changing these subsets of parameters tunes the free energy of the system in a way that is concordant with theoretical expectations. Finally, we show that the induction profiles and resulting free energies associated with pairwise double mutants can be predicted with quantitative accuracy given knowledge of the single mutants, providing an avenue for identifying and quantifying epistatic interactions. Summary We present a biophysical model of allosteric transcriptional regulation that directly links the location of a mutation within a repressor to the biophysical parameters that describe its behavior. We explore the phenotypic space of a repressor with mutations in either the inducer binding or DNA binding domains. Using the LacI repressor in E. coli , we make sharp, falsifiable predictions and use this framework to generate a null hypothesis for how double mutants behave given knowledge of the single mutants. Linking mutations to the parameters which govern the system allows for quantitative predictions of how the free energy of the system changes as a result, permitting coarse graining of high-dimensional data into a single-parameter description of the mutational consequences.
What are the thermodynamic costs of development? In this issue of Developmental Cell, Rodenfels et al. (2019) demonstrate that the high energetic cost of coordinated cell division that is regulated by phospho-signaling gives rise to a measurable periodicity in the heat dissipated during zebrafish embryogenesis.
It is tempting to believe that we now own the genome. The ability to read and re-write it at will has ushered in a stunning period in the history of science. Nonetheless, there is an Achilles heel exposed by all of the genomic data that has accrued: we still don't know how to interpret it. Many genes are subject to sophisticated programs of transcriptional regulation, mediated by DNA sequences that harbor binding sites for transcription factors which can up- or down-regulate gene expression depending upon environmental conditions. This gives rise to an input-output function describing how the level of expression depends upon the parameters of the regulated gene { for instance, on the number and type of binding sites in its regulatory sequence. In recent years, the ability to make precision measurements of expression, coupled with the ability to make increasingly sophisticated theoretical predictions, have enabled an explicit dialogue between theory and experiment that holds the promise of covering this genomic Achilles heel. The goal is to reach a predictive understanding of transcriptional regulation that makes it possible to calculate gene expression levels from DNA regulatory sequence. This review focuses on the canonical simple repression motif to ask how well the models that have been used to characterize it actually work. We consider a hierarchy of increasingly sophisticated experiments in which the minimal parameter set learned at one level is applied to make quantitative predictions at the next. We show that these careful quantitative dissections provide a template for a predictive understanding of the many more complex regulatory arrangements found across all domains of life.
Mutation is a critical mechanism by which evolution explores the functional landscape of proteins. Despite our ability to experimentally inflict mutations at will, it remains difficult to link sequence-level perturbations to systems-level responses. Here, we present a framework centered on measuring changes in the free energy of the system to link individual mutations in an allosteric transcriptional repressor to the parameters which govern its response. We find that the energetic effects of the mutations can be categorized into several classes which have characteristic curves as a function of the inducer concentration. We experimentally test these diagnostic predictions using the well-characterized LacI repressor of Escherichia coli, probing several mutations in the DNA binding and inducer binding domains. We find that the change in gene expression due to a point mutation can be captured by modifying only the model parameters that describe the respective domain of the wild-type protein. These parameters appear to be insulated, with mutations in the DNA binding domain altering only the DNA affinity and those in the inducer binding domain altering only the allosteric parameters. Changing these subsets of parameters tunes the free energy of the system in a way that is concordant with theoretical expectations. Finally, we show that the induction profiles and resulting free energies associated with pairwise double mutants can be predicted with quantitative accuracy given knowledge of the single mutants, providing an avenue for identifying and quantifying epistatic interactions.
Allosteric regulation is found across all domains of life, yet we still lack simple, predictive theories that directly link the experimentally tunable parameters of a system to its input-output response. To that end, we present a general theory of allosteric transcriptional regulation using the Monod-Wyman-Changeux model. We rigorously test this model using the ubiquitous simple repression motif in bacteria by first predicting the behavior of strains that span a large range of repressor copy numbers and DNA binding strengths and then constructing and measuring their response. Our model not only accurately captures the induction profiles of these strains but also enables us to derive analytic expressions for key properties such as the dynamic range and [EC50]. Finally, we derive an expression for the free energy of allosteric repressors which enables us to collapse our experimental data onto a single master curve that captures the diverse phenomenology of the induction profiles.
11 Allosteric molecules serve as regulators of cellular activity across all domains of life. We present a general 12 theory of allosteric transcriptional regulation that permits quantitative predictions for how physiological 13 responses are tuned to environmental stimuli. To test the model’s predictive power, we apply it to the 14 specific case of the ubiquitous simple repression motif in bacteria. We measure the fold-change in gene 15 expression at different inducer concentrations in a collection of strains that span a range of repressor 16 copy numbers and operator binding strengths. After inferring the inducer dissociation constants using 17 data from one of these strains, we show the broad reach of the model by predicting the induction profiles 18 of all other strains. Finally, we derive an expression for the free energy of allosteric transcription factors 19 which enables us to collapse the data from all of our experiments onto a single master curve, capturing 20 the diverse phenomenology of the induction profiles. 21
Allosteric molecules serve as regulators of cellular activity across all domains of life. We present a general theory of allosteric transcriptional regulation that permits quantitative predictions for how physiological responses are tuned to environmental stimuli. To test the model's predictive power, we apply it to the specific case of the ubiquitous simple repression motif in bacteria. We measure the fold-change in gene expression at different inducer concentrations in a collection of strains that span a range of repressor copy numbers and operator binding strengths. After inferring the inducer dissociation constants using data from one of these strains, we show the broad reach of the model by predicting the induction profiles of all other strains. Finally, we derive an expression for the free energy of allosteric transcription factors which enables us to collapse the data from all of our experiments onto a single master curve, capturing the diverse phenomenology of the induction profiles.
467 Manuel Razo-Mejia1,†, Stephanie L. Barnes1,†, Nathan M. Belliveau1,†, Griffin Chure1,†, 468 Tal Einav2,†, Rob Phillips1,3,∗ 469 Division of Biology and Biological Engineering, California Institute of Technology, Pasadena, United 470 States; Department of Physics, California Institute of Technology, Pasadena, United States; 471 Department of Applied Physics, California Institute of Technology, Pasadena, United States 472 † contributed equally 473