Organismal physiology is widely regulated by the molecular circadian clock, a feedback loop composed of protein complexes whose members are enriched in intrinsically disordered regions. These regions can mediate protein-protein interactions via SLiMs, but the contribution of these disordered regions to clock protein interactions had not been elucidated. To determine the functionality of these disordered regions, we applied a synthetic peptide microarray approach to the disordered clock protein FRQ in Neurospora crassa. We identified residues required for FRQ's interaction with its partner protein FRH, the mutation of which demonstrated FRH is necessary for persistent clock oscillations but not repression of transcriptional activity. Additionally, the microarray demonstrated an enrichment of FRH binding to FRQ peptides with a net positive charge. We found that positively charged residues occurred in significant "blocks" within the amino acid sequence of FRQ and that ablation of one of these blocks affected both core clock timing and physiological clock output. Finally, we found positive charge clusters were a commonly shared molecular feature in repressive circadian clock proteins. Overall, our study suggests a mechanistic purpose for positive charge blocks and yielded insights into repressive arm protein roles in clock function. Many clock proteins contain intrinsically disordered regions, but how these regions mediate protein interactions is poorly understood. Here, the authors identify charge blocks within a disordered clock protein that regulate circadian timing.
Sequence-specific activation by transcription factors is essential for gene regulation1,2. Key to this are activation domains, which often fall within disordered regions of transcription factors3,4 and recruit co-activators to initiate transcription5. These interactions are difficult to characterize via most experimental techniques because they are typically weak and transient6,7. Consequently, we know very little about whether these interactions are promiscuous or specific, the mechanisms of binding, and how these interactions tune the strength of gene activation. To address these questions, we developed a microfluidic platform for expression and purification of hundreds of activation domains in parallel followed by direct measurement of co-activator binding affinities (STAMMPPING, for Simultaneous Trapping of Affinity Measurements via a Microfluidic Protein-Protein INteraction Generator). By applying STAMMPPING to quantify direct interactions between eight co-activators and 204 human activation domains (>1,500 K ds), we provide the first quantitative map of these interactions and reveal 334 novel binding pairs. We find that the metazoan-specific co-activator P300 directly binds >100 activation domains, potentially explaining its widespread recruitment across the genome to influence transcriptional activation. Despite sharing similar molecular properties (e.g. enrichment of negative and hydrophobic residues), activation domains utilize distinct biophysical properties to recruit certain co-activator domains. Co-activator domain affinity and occupancy are well-predicted by analytical models that account for multivalency, and in vitro affinities quantitatively predict activation in cells with an ultrasensitive response. Not only do our results demonstrate the ability to measure affinities between even weak protein-protein interactions in high throughput, but they also provide a necessary resource of over 1,500 activation domain/co-activator affinities which lays the foundation for understanding the molecular basis of transcriptional activation.
ABSTRACTIntrinsically disordered regions (IDRs) are ubiquitous across all domains of life and play a range of functional roles. While folded domains are generally well-described by a single 3D structure, IDRs exist in a collection of interconverting states known as an ensemble. This structural heterogeneity means IDRs are largely absent from the PDB, contributing to a lack of computational approaches to predict ensemble conformational properties from sequence. Here we combine rational sequence design, large-scale molecular simulations, and deep learning to develop ALBATROSS, a deep learning model for predicting IDR ensemble dimensions from sequence. ALBATROSS enables the instantaneous prediction of ensemble average properties at proteome-wide scale. ALBATROSS is lightweight, easy-to-use, and accessible as both a locally installable software package and a point-and-click style interface in the cloud. We first demonstrate the applicability of our predictors by examining the generalizability of sequence-ensemble relationships in IDRs. Then, we leverage the high-throughput nature of ALBATROSS to characterize emergent biophysical behavior of IDRs within and between proteomes.Update from previous versionThis preprint reports an updated version of the ALBATROSS network weights trained on simulations of over 42,000 sequences.In addition, we provide new colab notebooks that enable proteome-wide IDR prediction and annotation in minutes.All conclusions and observations made in versions 1 and 2 of this manuscript remain true and robust.
Intrinsically disordered proteins and protein regions (collectively IDRs) are critical in numerous cellular processes. To understand how IDRs facilitate function, we need tools to accurately and rapidly identify them from sequence. While many methods for disorder prediction exist, we are currently limited by throughput and accuracy for evolutionary scale analyses. To bridge this gap, we developed metapredict V3, an updated version of our disorder predictor that enables evolutionary-scale disorder prediction. Metapredict V3 enables proteome-scale prediction with state-of-the-art accuracy in seconds and was developed with a focus on usability. It is distributed as a web server, Python software package, command-line interface, and Google Colab notebook. Here, we leverage the accuracy and throughput of metapredict V3 to predict disorder for over 20,000 proteomes to evaluate the prevalence of disorder across the kingdoms of life. ### Competing Interest Statement ASH is on the scientific advisory board of Prose Foods. All other authors declare no competing interests.
Human gene expression is regulated by over two thousand transcription factors and chromatin regulator proteins. Activation domains (ADs) within these proteins recruit shared co-activators and Pol II to genes to activate transcription. However, for many human ADs we do not know which co-activators they recruit, if ADs bind single co-activators specifically or many promiscuously, and why some ADs are stronger at activating gene expression than others. In previous work, we have measured the effect on gene expression for thousands of human ADs and mutants inside cells. Here, we developed a microfluidic platform for expression and purification of hundreds of these ADs in parallel followed by direct measurement of AD/co-activator binding affinities (STAMMPPING, for Simultaneous Trapping of Affinity Measurements via a Microfluidic Protein-Protein INteraction Generator). STAMMPPING can reliably quantify interaction strengths ranging across two orders of magnitude (Kds from 0.5 to 50 μM). To date, we used STAMMPPING to measure over 3000 interactions between 204 ADs and 15 co-activators, identifying both specific and promiscuous interactions. Affinities between ADs and individual domains of the co-activator P300 identified ADs that bind multiple domains of the same co-activator, suggesting longer lived total interactions can be achieved with individual short-lived interactions. Complementary mutagenesis of ADs revealed that decreasing their strength of binding to co-activators proportionally decreases their ability to activate gene expression in cells. Our systematic quantification of AD/co-activator interactions advances our understanding of how human ADs activate gene expression. Moreover, direct affinity data produced here will be a rich resource for mapping gene regulatory networks and building interpretable predictive models of AD/co-activator binding.
Biomolecular condensation underlies the biogenesis of an expanding array of membraneless assemblies, including stress granules (SGs), which form under a variety of cellular stresses. Advances have been made in understanding the molecular grammar of a few scaffold proteins that make up these phases, but how the partitioning of hundreds of SG proteins is regulated remains largely unresolved. While investigating the rules that govern the condensation of ataxin-2, an SG protein implicated in neurodegenerative disease, we unexpectedly identified a short 14 aa sequence that acts as a condensation switch and is conserved across the eukaryote lineage. We identify poly(A)-binding proteins as unconventional RNA-dependent chaperones that control this regulatory switch. Our results uncover a hierarchy of cis and trans interactions that fine-tune ataxin-2 condensation and reveal an unexpected molecular function for ancient poly(A)-binding proteins as regulators of biomolecular condensate proteins. These findings may inspire approaches to therapeutically target aberrant phases in disease.
Many important biological processes are regulated by the circadian molecular clock, a negative feedback loop comprised of disordered proteins. The lack of a fixed conformation means that identifying interaction regions on core clock proteins is untenable using typical biophysical methods. Since disordered proteins often interact through short linear binding motifs (SLiMs), we decided to take a divide-and-conquer approach leveraging synthetic peptide technology. Our method termed Linear mOtif disCovery using rATional dEsign (LOCATE) takes a disordered protein sequence and divides it into a set of short overlapping peptides that can be synthesized and printed in a microarray format.
SummaryThe circadian clock times cellular processes to the day/night cycle via a Transcription-Translation negative Feedback Loop (TTFL). However, a mechanistic understanding of the negative arm in both the timing of the TTFL and its control of output is lacking. We posited that the formation of negative-arm protein complexes was fundamental to clock regulation stemming from the negative arm. Using a modified peptide microarray approach termed Linear motif discovery using rational design (LOCATE), we characterized the interaction of the disordered negative-arm clock protein FREQUENCY to its partner protein FREQUENCY-Interacting RNA helicase. LOCATE identified a specific Short Linear Motif (SLiM) and interaction “hotspot” as well as positively charged “islands” that mediate electrostatic interactions, suggesting a model where negative arm proteins form a “fuzzy” complex essential for clock timing and robustness. Further analysis revealed that the positively charged islands were an evolutionarily conserved feature in higher eukaryotes and contributed to proper clock function.
Organismal physiology is widely regulated by the circadian clock, a molecular circuit composed of a Transcription-Translation Feedback Loop 1,2. Protein components of the molecular clock are enriched in intrinsically disordered regions, inherently flexible regions that interact with other proteins via short linear binding motifs (SLiMs) 3–5. SLiM-driven interactions contribute to circadian timing and the circadian regulation of the cell. However, the mechanism that allows the formation of dynamic clock complexes remains unclear as structural analysis of these protein-protein interactions has been limited due to inherent protein disorder. Here, we apply a synthetic peptide microarray approach to demonstrate that the core clock forms a fuzzy complex to support circadian robustness 6,7. We found positively charged islands on the clock protein FREQUENCY (FRQ) drove a multi-valent interaction between FRQ and its partner FRQ-interacting RNA Helicase (FRH) that enabled clock robustness rather than the previously-reported feedback 8. We found these positively charged islands were a conserved molecular feature throughout clocks in fungi, insects, and mammals, and may enable the formation of fuzzy complexes. This study constitutes the first mechanistic reason for the uniquely-broad conservation of intrinsic disorder in circadian negative-arm proteins and will aid in the development of the molecular model of clock protein interactions. Furthermore, we anticipate the application of synthetic peptide microarrays to study disordered clock proteins and will be useful in characterizing sites of interaction for clock-specific drug discovery 9.
Intrinsically disordered proteins and protein regions make up a substantial fraction of many proteomes in which they play a wide variety of essential roles. A critical first step in understanding the role of disordered protein regions in biological function is to identify those disordered regions correctly. Computational methods for disorder prediction have emerged as a core set of tools to guide experiments, interpret results, and develop hypotheses. Given the multiple different predictors available, consensus scores have emerged as a popular approach to mitigate biases or limitations of any single method. Consensus scores integrate the outcome of multiple independent disorder predictors and provide a per-residue value that reflects the number of tools that predict a residue to be disordered. Although consensus scores help mitigate the inherent problems of using any single disorder predictor, they are computationally expensive to generate. They also necessitate the installation of multiple different software tools, which can be prohibitively difficult. To address this challenge, we developed a deep-learning-based predictor of consensus disorder scores. Our predictor, metapredict, utilizes a bidirectional recurrent neural network trained on the consensus disorder scores from 12 proteomes. By benchmarking metapredict using two orthogonal approaches, we found that metapredict is among the most accurate disorder predictors currently available. Metapredict is also remarkably fast, enabling proteome-scale disorder prediction in minutes. Importantly, metapredict is a fully open source and is distributed as a Python package, a collection of command-line tools, and a web server, maximizing the potential practical utility of the predictor. We believe metapredict offers a convenient, accessible, accurate, and high-performance predictor for single-proteins and proteomes alike.
The SARS-CoV-2 nucleocapsid (N) protein is an abundant RNA binding protein critical for viral genome packaging, yet the molecular details that underlie this process are poorly understood. Here we combine single-molecule spectroscopy with all-atom simulations to uncover the molecular details that contribute to N protein function. N protein contains three dynamic disordered regions that house putative transiently-helical binding motifs. The two folded domains interact minimally such that full-length N protein is a flexible and multivalent RNA binding protein. N protein also undergoes liquid-liquid phase separation when mixed with RNA, and polymer theory predicts that the same multivalent interactions that drive phase separation also engender RNA compaction. We offer a simple symmetry-breaking model that provides a plausible route through which single-genome condensation preferentially occurs over phase separation, suggesting that phase separation offers a convenient macroscopic readout of a key nanoscopic interaction.
In immature oocytes, Balbiani bodies are conserved membraneless condensates implicated in oocyte polarization, the organization of mitochondria, and long-term organelle and RNA storage. In Xenopus laevis, Balbiani body assembly is mediated by the protein Velo1. Velo1 contains an N-terminal prion-like domain (PLD) that is essential for Balbiani body formation. PLDs have emerged as a class of intrinsically disordered regions that can undergo various different types of intracellular phase transitions and are often associated with dynamic, liquid-like condensates. Intriguingly, the Velo1 PLD forms solid-like assemblies. Here we sought to understand why Velo1 phase behavior appears to be biophysically distinct from that of other PLD-containing proteins. Through bioinformatic analysis and coarse-grained simulations, we predict that the clustering of aromatic residues and the amino acid composition of residues between aromatics can influence condensate material properties, organization, and the driving forces for assembly. To test our predictions, we redesigned the Velo1 PLD to test the impact of targeted sequence changes in vivo. We found that the Velo1 design with evenly spaced aromatic residues shows rapid internal dynamics, as probed by fluorescent recovery after photobleaching, even when recruited into Balbiani bodies. Our results suggest that Velo1 might have been selected in evolution for distinctly clustered aromatic residues to maintain the structure of Balbiani bodies in long-lived oocytes. In general, our work identifies several tunable parameters that can be used to augment the condensate material state, offering a road map for the design of synthetic condensates.
The rise of high-throughput experiments has transformed how scientists approach biological questions. The ubiquity of large-scale assays that can test thousands of samples in a day has necessitated the development of new computational approaches to interpret this data. Among these tools, machine learning approaches are increasingly being utilized due to their ability to infer complex nonlinear patterns from high-dimensional data. Despite their effectiveness, machine learning (and in particular deep learning) approaches are not always accessible or easy to implement for those with limited computational expertise. Here we present PARROT, a general framework for training and applying deep learning-based predictors on large protein datasets. Using an internal recurrent neural network architecture, PARROT is capable of tackling both classification and regression tasks while only requiring raw protein sequences as input. We showcase the potential uses of PARROT on three diverse machine learning tasks: predicting phosphorylation sites, predicting transcriptional activation function of peptides generated by high-throughput reporter assays, and predicting the fibrillization propensity of amyloid beta with data generated by deep mutational scanning. Through these examples, we demonstrate that PARROT is easy to use, performs comparably to state-of-the-art computational tools, and is applicable for a wide array of biological problems.
L-Tyrosine is an essential aromatic amino acid required for the synthesis of proteins and a diverse array of plant natural products; however, little is known on how the levels of tyrosine are controlled in planta and linked to overall growth and development. Most plants synthesize tyrosine by TyrA arogenate dehydrogenases, which are strongly feedback-inhibited by tyrosine and encoded by TyrA1 and TyrA2 genes in Arabidopsis thaliana. While TyrA enzymes have been extensively characterized at biochemical levels, their in planta functions remain uncertain. Here we found that TyrA1 suppression reduces seed yield due to impaired anther dehiscence, whereas TyrA2 knockout leads to slow growth with reticulate leaves. The tyra2 mutant phenotypes were exacerbated by TyrA1 suppression and rescued by the expression of TyrA2, TyrA1 or tyrosine feeding. Low-light conditions synchronized the tyra2 and wild-type growth, and ameliorated the tyra2 leaf reticulation. After shifting to normal light, tyra2 transiently decreased tyrosine and subsequently increased aspartate before the appearance of the leaf phenotypes. Overexpression of the deregulated TyrA enzymes led to hyper-accumulation of tyrosine, which was also accompanied by elevated aspartate and reticulate leaves. These results revealed that TyrA1 and TyrA2 have distinct and overlapping functions in flower and leaf development, respectively, and that imbalance of tyrosine, caused by altered TyrA activity and regulation, impacts growth and development of Arabidopsis. The findings provide critical bases for improving the production of tyrosine and its derived natural products, and further elucidating the coordinated metabolic and physiological processes to maintain tyrosine levels in plants.