Abstract Enhancers coordinate gene expression in response to developmental and environmental cues. Because plant enhancers lack the readily detectable molecular hallmarks of animal enhancers, their systematic functional characterization has yet to be accomplished. Here, we characterize the species- and condition-specific enhancer activity of over 350,000 sequences derived from accessible chromatin regions of the Arabidopsis , tomato, maize, and sorghum genomes. Enabled by the massive scale of the data, we developed plantGREP, a deep learning model that predicts enhancer strength and identifies the underlying functional sequence motifs. We apply plantGREP to evolve strong constitutive as well as species- and condition-specific enhancers, and to locate regions with enhancer activity upstream of developmental genes in crop genomes. These results should facilitate the targeted editing of enhancers in crop genomes and the design of cell-type-specific plant enhancers.
Deep mutational scanning couples a protein's activity to DNA sequencing for high throughput assessment of the effects of all single amino acid substitutions, but it largely uses indirect assays, like growth, as proxy for protein activity. Here, we covalently link variant proteins in vivo to an RNA barcode by fusing them to E. coli tRNA (m5U54) methyltransferase TrmA (E358Q), which forms a covalent bond with a tRNA stem-loop. Following cell lysis, variant proteins are separated in vitro according to their biochemical properties and identified by their barcodes. We use this method, Dosa, to analyze a large pool of FLAG epitope variants for binding to an anti-FLAG antibody, to profile the cleavage preferences of variants of enteropeptidase and human rhinovirus 3C protease, and to measure the solubility of several hundred Aβ(1-42) variants. This method should be amenable to numerous biochemical assays with proteins produced in E. coli or mammalian cells.
Insulators are cis-regulatory elements that separate transcriptional units, whereas silencers are elements that repress transcription regardless of their position. In plants, these elements remain largely uncharacterized. Here, we use the massively parallel reporter assay Plant STARR-seq with short fragments of 8 large insulators to identify more than 100 fragments that block enhancer activity. The short fragments can be combined to generate more powerful insulators that abolish the capacity of the strong viral 35S enhancer to activate the 35S minimal promoter. Unexpectedly, when tested upstream of weak enhancers, these fragments act as silencers and repress transcription. Thus, these elements are capable of insulating or repressing transcription, depending on the regulatory context. We validate our findings in stable transgenic Arabidopsis thaliana, maize (Zea mays), and rice (Oryza sativa) plants. The short elements identified here should be useful building blocks for plant biotechnology.
The 3' end of a gene, often called a terminator, modulates mRNA stability, localization, translation, and polyadenylation. Here, we adapted Plant STARR-seq, a massively parallel reporter assay, to measure the activity of over 50,000 terminators from the plants Arabidopsis thaliana and Zea mays. We characterize thousands of plant terminators, including many that outperform bacterial terminators commonly used in plants. Terminator activity is species-specific, differing in tobacco leaf and maize protoplast assays. While recapitulating known biology, our results reveal the relative contributions of polyadenylation motifs to terminator strength. We built a computational model to predict terminator strength and used it to conduct in silico evolution that generated optimized synthetic terminators. Additionally, we discover alternative polyadenylation sites across tens of thousands of terminators; however, the strongest terminators tend to have a dominant cleavage site. Our results establish features of plant terminator function and identify strong naturally occurring and synthetic terminators.
Enhancers are cis-regulatory elements that shape gene expression in response to numerous developmental and environmental cues. In animals, several models have been proposed to explain how enhancers integrate the activity of multiple transcription factors. However, it remains largely unclear how plant enhancers integrate transcription factor activity. Here, we use Plant STARR-seq to characterize 3 light-responsive plant enhancers-AB80, Cab-1, and rbcS-E9-derived from genes associated with photosynthesis. Saturation mutagenesis revealed mutations, many of which clustered in short regions, that strongly reduced enhancer activity in the light, in the dark, or in both conditions. When tested in the light, these mutation-sensitive regions did not function on their own; rather, cooperative interactions with other such regions were required for full activity. Epistatic interactions occurred between mutations in adjacent mutation-sensitive regions, and the spacing and order of mutation-sensitive regions in synthetic enhancers affected enhancer activity. In contrast, when tested in the dark, mutation-sensitive regions acted independently and additively in conferring enhancer activity. Taken together, this work demonstrates that plant enhancers show evidence for both cooperative and additive interactions among their functional elements. This knowledge can be harnessed to design strong, condition-specific synthetic enhancers.
Intron splicing is a key regulatory step in gene expression in eukaryotes. Three sequence elements required for splicing-5 ' and 3 ' splice sites and a branchpoint-are especially well-characterized in Saccharomyces cerevisiae, but our understanding of additional intron features that impact splicing in this organism is incomplete, due largely to its small number of introns. To overcome this limitation, we constructed a library in S. cerevisiae of random 50-nt (N50) elements individually inserted into the intron of a reporter gene and quantified canonical splicing and the use of cryptic splice sites by sequencing analysis. More than 70% of approximately 140,000 N50 elements reduced splicing by at least 20%. N50 features, including higher GC content, presence of GU repeats, and stronger predicted secondary structure of its pre-mRNA, correlated with reduced splicing efficiency. A likely basis for the reduced splicing of such a large proportion of variants is the formation of RNA structures that pair N50 bases-such as the GU repeats-with other bases specifically within the reporter pre-mRNA analyzed. However, multiple models were unable to explain more than a small fraction of the variance in splicing efficiency across the library, suggesting that complex nonlinear interactions in RNA structures are not accurately captured by RNA structure prediction methods. Our results imply that the specific context of a pre-mRNA may determine the bases allowable in an intron to prevent secondary structures that reduce splicing. This large data set can serve as a resource for further exploration of splicing mechanisms.
Over the last three decades, human genetics has gone from dissecting high-penetrance Mendelian diseases to discovering the vast and complex genetic etiology of common human diseases. In tackling this complexity, scientists have discovered the importance of numerous genetic processes – most notably functional regulatory elements – in the development and progression of these diseases. Simultaneously, scientists have increasingly used multiplex assays of variant effect to systematically phenotype the cellular consequences of millions of genetic variants. In this article, we argue that the context of genetic variants – at all scales, from other genetic variants and gene regulation to cell biology to organismal environment – are critical components of how we can employ genomics to interpret these variants, and ultimately treat these diseases. We describe approaches to extend existing experimental assays and computational approaches to examine and quantify the importance of this context, including through causal analytic approaches. Having a unified understanding of the molecular, physiological, and environmental processes governing the interpretation of genetic variants is sorely needed for the field, and this perspective argues for feasible approaches by which the combined interpretation of cellular, animal, and epidemiological data can yield that knowledge.
The 3’ end of a gene, often called a terminator, modulates mRNA stability, localization, translation, and polyadenylation. Here, we adapted Plant STARR-seq, a massively parallel reporter assay, to measure the activity of over 50,000 terminators from the plants Arabidopsis thaliana and Zea mays. We characterize thousands of plant terminators, including many that outperform bacterial terminators commonly used in plants. Terminator activity is speciesspecific, differing in tobacco leaf and maize protoplast assays. While recapitulating known biology, our results reveal the relative contributions of polyadenylation motifs to terminator strength. We built a computational model to predict terminator strength and used it to conduct in silico evolution that generated optimized synthetic terminators. Additionally, we discover alternative polyadenylation sites across tens of thousands of terminators; however, the strongest terminators tend to have a dominant cleavage site. Our results establish features of plant terminator function and identify strong naturally occurring and synthetic terminators.
Massively parallel measurements of dominant-negative inhibition by protein fragments have been used to map protein interaction sites and discover peptide inhibitors. However, the underlying principles governing fragment-based inhibition have thus far remained unclear. Here, we adapted a high-throughput inhibitory fragment assay for use in Escherichia coli, applying it to a set of 10 essential proteins. This approach yielded single amino acid resolution maps of inhibitory activity, with peaks localized to functionally important interaction sites, including oligomerization interfaces and folding contacts. Leveraging these data, we performed a systematic analysis to uncover principles of fragment-based inhibition. We determined a robust negative correlation between susceptibility to inhibition and cellular protein concentration, demonstrating that inhibitory fragments likely act primarily by titrating native protein interactions. We also characterized a series of trade-offs related to fragment length, showing that shorter peptides allow higher-resolution mapping but suffer from lower inhibitory activity. We employed an unsupervised statistical analysis to show that the inhibitory activities of protein fragments are largely driven not by generic properties such as charge, hydrophobicity, and secondary structure, but by the more specific characteristics of their bespoke macromolecular interactions. Overall, this work demonstrates fundamental characteristics of inhibitory protein fragment function and provides a foundation for understanding and controlling protein interactions in vivo.
ABSTRACT Enhancers are cis -regulatory elements that shape gene expression in response to numerous developmental and environmental cues. In animals, several models have been proposed to explain how enhancers integrate the activity of multiple transcription factors. However, it remains largely unknown how plant enhancers integrate transcription factor activity. Here, we use Plant STARR-seq to characterize three light-responsive plant enhancers— AB80 , Cab-1 , and rbcS-E9 —derived from genes active in photosynthesis. Saturation mutagenesis reveals mutations, many of which cluster in short regions, that strongly reduce enhancer activity in the light, in the dark or in both conditions. When tested in the light, these mutation-sensitive regions do not function on their own; rather, cooperative interactions with other such regions are required for full activity. Epistatic interactions occur between mutations in adjacent mutation-sensitive regions, and the spacing and order of mutation-sensitive regions in synthetic enhancers affects enhancer activity. In contrast, when tested in the dark, mutation-sensitive regions act independently and additively in conferring enhancer activity. Taken together, this work demonstrates that plant enhancers show evidence for both cooperative and additive interactions among their functional elements. This knowledge can be harnessed to design strong, condition-specific synthetic enhancers.
DNA sequencing has led to the discovery of millions of mutations that change the encoded protein sequences, but the impact of nearly all of these mutations on protein function is unknown. We addressed this scarcity of functional data by developing Miro, a proteomic technology that uses mistranslation to introduce amino acid substitutions and biochemical assays to quantify functional differences of thousands of protein variants by mass spectrometry. We apply this technology to the proteome of yeast to reveal amino acid substitutions that impact protein structure, ligand binding, protein-protein interactions, protein post-translational modifications, and protein thermal stability. Adapting Miro to human cells will provide a means to efficiently accelerate our mechanistic interpretation of genomic mutations to predict disease risk.
Antibiotic resistance is a growing threat to public health, making the development of antibiotics of critical importance. One promising class of potential new antibiotics are ribosomally synthesized and post-translationally modified peptides (RiPPs), which include klebsidin, a lasso peptide from Klebsiella pneumoniae that inhibits certain bacterial RNA polymerases. We develop a high-throughput assay based on growth inhibition of Escherichia coli to analyze the mutational tolerance of klebsidin. We transform a library of klebsidin variants into E. coli and use next-generation DNA sequencing to count the frequency of each variant before and after its expression, thereby generating functional scores for 320 of 361 single amino acid changes. We identify multiple positions in the macrocyclic ring and the C-terminal tail region of klebsidin that are intolerant to mutation, as well as positions in the loop region that are highly tolerant to mutation. Characterization of selected peptide variants scored as active reveals that each adopts a threaded lasso conformation; active loop variants applied extracellularly as peptides slow the growth of E. coli and K. pneumoniae. We generate an E. coli strain with a mutation in RNA polymerase that confers resistance to klebsidin and similarly carry out a selection with the klebsidin library. We identify a single variant, klebsidin F9Y, that maintains activity against the resistant E. coli when expressed intracellularly. This finding supports the utility of this method and suggests that comprehensive mutational analysis of lasso peptides can identify unique and potentially improved variants.
Background: The 3 ' untranslated region (UTR) plays critical roles in determining the level of gene expression through effects on activities such as mRNA stability and translation. Functional elements within this region have largely been identified through analyses of native genes, which contain multiple co-evolved sequence features. Results: To explore the effects of 3 ' UTR sequence elements outside of native sequence contexts, we analyze hundreds of thousands of random 50-mers inserted into the 3 ' UTR of a reporter gene in the yeast Saccharomyces cerevisiae. We determine relative protein expression levels from the fitness of transformants in a growth selection. We find that the consensus 3 ' UTR efficiency element significantly boosts expression, independent of sequence context; on the other hand, the consensus positioning element has only a small effect on expression. Some sequence motifs that are binding sites for Puf proteins substantially increase expression in the library, despite these proteins generally being associated with post-transcriptional downregulation of native mRNAs. Our measurements also allow a systematic examination of the effects of point mutations within efficiency element motifs across diverse sequence backgrounds. These mutational scans reveal the relative in vivo importance of individual bases in the efficiency element, which likely reflects their roles in binding the Hrp1 protein involved in cleavage and polyadenylation. Conclusions: The regulatory effects of some 3 ' UTR sequence features, like the efficiency element, are consistent regardless of sequence context. In contrast, the consequences of other 3 ' UTR features appear to be strongly dependent on their evolved context within native genes.
SUMMARY:Multiplexed assays of variant effect (MAVEs) are capable of experimentally testing all possible single nucleotide or amino acid variants in selected genomic regions, generating 'variant effect maps', which provide biochemical insight and functional evidence to enable more rapid and accurate clinical interpretation of human variation. Because the international community applying MAVE approaches is growing rapidly, we developed the online MaveRegistry platform to catalyze collaboration, reduce redundant efforts, allow stakeholders to nominate targets and enable tracking and sharing of progress on ongoing MAVE projects. AVAILABILITY AND IMPLEMENTATION:MaveRegistry service: https://registry.varianteffect.org. MaveRegistry source code: https://github.com/kvnkuang/maveregistry-front-end.
The ability to design a protein to bind specifically to a target RNA enables numerous applications, with the modular architecture of the PUF domain lending itself to new RNA-binding specificities. For each repeat of the Pumilio-1 PUF domain, we generate a library that contains the 8,000 possible combinations of amino acid substitutions at residues critical for RNA contact. We carry out yeast three-hybrid selections with each library against the RNA recognition sequence for Pumilio-1, with any possible base present at the position recognized by the randomized repeat. We use sequencing to score the binding of each variant, identifying many variants with highly repeat-specific interactions. From these data, we generate an RNA binding code specific to each repeat and base. We use this code to design PUF domains against 16 RNAs, and find that some of these domains recognize RNAs with two, three or four changes from the wild type sequence.
Targeted engineering of plant gene expression holds great promise for ensuring food security and for producing biopharmaceuticals in plants. However, this engineering requires thorough knowledge of cis-regulatory elements to precisely control either endogenous or introduced genes. To generate this knowledge, we used a massively parallel reporter assay to measure the activity of nearly complete sets of promoters from Arabidopsis, maize and sorghum. We demonstrate that core promoter elements-notably the TATA box-as well as promoter GC content and promoter-proximal transcription factor binding sites influence promoter strength. By performing the experiments in two assay systems, leaves of the dicot tobacco and protoplasts of the monocot maize, we detect species-specific differences in the contributions of GC content and transcription factors to promoter strength. Using these observations, we built computational models to predict promoter strength in both assay systems, allowing us to design highly active promoters comparable in activity to the viral 35S minimal promoter. Our results establish a promising experimental approach to optimize native promoter elements and generate synthetic ones with desirable features.
The scarcity of accessible sites that are dynamic or cell type-specific in plants may be due in part to tissue heterogeneity in bulk studies. To assess the effects of tissue heterogeneity, we apply single-cell ATAC-seq to Arabidopsis thaliana roots and identify thousands of differentially accessible sites, sufficient to resolve all major cell types of the root. We find that the entirety of a cell's regulatory landscape and its transcriptome independently capture cell type identity. We leverage this shared information on cell identity to integrate accessibility and transcriptome data to characterize developmental progression, endoreduplication and cell division. We further use the combined data to characterize cell type-specific motif enrichments of transcription factor families and link the expression of family members to changing accessibility at specific loci, resolving direct and indirect effects that shape expression. Our approach provides an analytical framework to infer the gene regulatory networks that execute plant development.