The variability of proteins at the sequence level creates an enormous potential for proteome complexity. Exploring the depths and limits of this complexity is an ongoing goal in biology. Here, we systematically survey human and plant high-throughput bottom-up native proteomics data for protein truncation variants, where substantial regions of the full-length protein are missing from an observed protein product. In humans, Arabidopsis, and the green alga Chlamydomonas, approximately one percent of observed proteins show a short form, which we can assign by comparison to RNA isoforms as either likely deriving from transcript-directed processes or limited proteolysis. While some detected protein fragments align with known splice forms and protein cleavage events, multiple examples are previously undescribed, such as our observation of fibrocystin proteolysis and nuclear translocation in a green alga. We find that truncations occur almost entirely between structured protein domains, even when short forms are derived from transcript variants. Intriguingly, multiple endogenous protein truncations of phase-separating translational proteins resemble cleaved proteoforms produced by enteroviruses during infection. Some truncated proteins are also observed in both humans and plants, suggesting that they date to the last eukaryotic common ancestor. Finally, we describe novel proteoform-specific protein complexes, where the loss of a domain may accompany complex formation.
Cryo-electron microscopy is traditionally applied to samples purified to near homogeneity as current reconstruction algorithms are unable to handle heterogeneous mixtures of structures from many macromolecular complexes. We extend on long established methods and demonstrate that relating two-dimensional projection images by their common lines in a graphical framework is sufficient for partitioning distinct protein and multiprotein complexes within the same data set. Using this approach, we first group a large set of synthetic reprojections from 35 unique macromolecular structures ranging from ∼30 – 3000 kDa into individual homogenous classes. We then apply our algorithm on cryo-EM data collected from a mixture of five protein complexes and use existing reconstruction methods to solve multiple three-dimensional structures ab initio . Incorporating methods to sort cryo-EM data from heterogeneous mixtures will alleviate the need for stringent purification and pave the way toward investigation of samples containing many unique structures.
RNA-binding proteins (RBPs) play essential roles in biology and are frequently associated with human disease. Although recent studies have systematically identified individual RNA-binding proteins, their higher-order assembly into ribonucleoprotein (RNP) complexes has not been systematically investigated. Here, we describe a proteomics method for systematic identification of RNP complexes in human cells. We identify 1,428 protein complexes that associate with RNA, indicating that more than 20% of known human protein complexes contain RNA. To explore the role of RNA in the assembly of each complex, we identify complexes that dissociate, change composition, or form stable protein-only complexes in the absence of RNA. We use our method to systematically identify cell-type-specific RNA-associated proteins in mouse embryonic stem cells and finally, distribute our resource, rna.MAP, in an easy-to-use online interface (rna.proteincomplexes.org). Our system thus provides a methodology for explorations across human tissues, disease states, and throughout all domains of life.
Multi-protein complexes are necessary for nearly all cellular processes, and understanding their structure is required for elucidating their function. Current high-resolution strategies in structural biology are effective, but lag behind other fields (e.g. genomics and proteomics) due to their reliance on purified samples rather than characterizing heterogeneous mixtures. Here, we present a method combining single particle analysis by electron microscopy with protein identification by mass spectrometry to structurally characterize macromolecular complexes from extracts of human cells. We obtain three-dimensional structures of native proteasomes directly from ab initio classification of a heterogeneous mixture of protein complexes. In addition, we find an ~1 MDa size structure of unknown composition and reference our proteomics data to suggest possible identities. Our study shows the power of using a shotgun approach to electron microscopy (shotgun EM) when coupled with mass spectrometry as a tool to uncover the structures of macromolecular machines in parallel.
Recent mass spectrometry maps of the human interactome independently support the existence of a large multiprotein complex, dubbed "Commander.'' Broadly conserved across animals and ubiquitously expressed in nearly every human cell type examined thus far, Commander likely plays a fundamental cellular function, akin to other ubiquitous machines involved in expression, proteostasis, and trafficking. Experiments on individual subunits support roles in endosomal protein sorting, including the trafficking of Notch proteins, copper transporters, and lipoprotein receptors. Commander is critical for vertebrate embryogenesis, and defects in the complex and its interaction partners disrupt craniofacial, brain, and heart development. Here, we review the synergy between large-scale proteomic efforts and focused studies in the discovery of Commander, describe its composition, structure, and function, and discuss how it illustrates the power of systems biology. Based on 3D modeling and biochemical data, we draw strong parallels between Commander and the retromer cargo-recognition complex, laying a foundation for future research into Commander's role in human developmental disorders.
How helicase families with a conserved catalytic 'helicase core' evolved to function on varied RNA and DNA substrates by diverse mechanisms remains unclear. Here, we used the helicase core of Mss116, a DEAD-box protein that utilizes ATP to locally unwind dsRNA, to investigate helicase specificity and mechanism. Previously, we found that the two RecA-like domains of the helicase core of Mss116 are in an extended 'open state' in the absence of substrates and recognize ATP and duplex RNA in a modular manner. Upon formation of a compact 'closed state' containing an ATPase active site, conserved motifs in the first domain promote the nonprocessive unwinding of short duplex substrates bound to the second domain by excluding one RNA strand and bending the other. In the present work, we define the molecular basis for the specificity of DEAD-box proteins. However, we also find that Mss116 has ambiguous substrate unwinding properties and interacts with a variety of NTPs and nucleic acids. The efficiency of unwinding correlates with the stability of the closed-state helicase core, a complex with nucleotide and nucleic acid that forms as duplexes are unwound. Crystal structures reveal that core stability is modulated by family-specific interactions that favor certain substrates. This suggests how present-day helicases diversified from an ancestral core with broad specificity by retaining core closure as a common catalytic mechanism while optimizing substrate-binding interactions for different cellular functions.
YbeA is a 3-methylpseudoridine methyltransferase from Escherichia coli that forms a stable homodimer in solution. It is one of the deeply trefoil 31 knotted proteins, of which the knot encompasses the C-terminal helix that threads through a long loop. Recent studies on the knotted protein folding pathways using YbeA have suggested that the protein knot remains present under chemically denaturing conditions. Here, we report (1)H, (13)C and (15)N chemical shift assignments for urea-denatured YbeA, which will serve as the basis for further structural characterisations using solution state NMR spectroscopy with paramagnetic spin labeled and partial alignment media.
How different helicase families with a conserved catalytic 'helicase core' evolved to function on varied RNA and DNA substrates by diverse mechanisms remains unclear. In this study, we used Mss116, a yeast DEAD-box protein that utilizes ATP to locally unwind dsRNA, to investigate helicase specificity and mechanism. Our results define the molecular basis for the substrate specificity of a DEAD-box protein. Additionally, they show that Mss116 has ambiguous substrate-binding properties and interacts with all four NTPs and both RNA and DNA. The efficiency of unwinding correlates with the stability of the 'closed-state' helicase core, a complex with nucleotide and nucleic acid that forms as duplexes are unwound. Crystal structures reveal that core stability is modulated by family-specific interactions that favor certain substrates. This suggests how present-day helicases diversified from an ancestral core with broad specificity by retaining core closure as a common catalytic mechanism while optimizing substrate-binding interactions for different cellular functions.
The Neurospora crassa mitochondrial tyrosyl-tRNA synthetase (mtTyrRS; CYT-18 protein) evolved a new function as a group I intron splicing factor by acquiring the ability to bind group I intron RNAs and stabilize their catalytically active RNA structure. Previous studies showed: (i) CYT-18 binds group I introns by using both its N-terminal catalytic domain and flexibly attached C-terminal anticodon-binding domain (CTD); and (ii) the catalytic domain binds group I introns specifically via multiple structural adaptations that occurred during or after the divergence of Peziomycotina and Saccharomycotina. However, the function of the CTD and how it contributed to the evolution of splicing activity have been unclear. Here, small angle X-ray scattering analysis of CYT-18 shows that both CTDs of the homodimeric protein extend outward from the catalytic domain, but move inward to bind opposite ends of a group I intron RNA. Biochemical assays show that the isolated CTD of CYT-18 binds RNAs non-specifically, possibly contributing to its interaction with the structurally different ends of the intron RNA. Finally, we find that the yeast mtTyrRS, which diverged from Pezizomycotina fungal mtTyrRSs prior to the evolution of splicing activity, binds group I intron and other RNAs non-specifically via its CTD, but lacks further adaptations needed for group I intron splicing. Our results suggest a scenario of constructive neutral (i.e., pre-adaptive) evolution in which an initial non-specific interaction between the CTD of an ancestral fungal mtTyrRS and a self-splicing group I intron was "fixed" by an intron RNA mutation that resulted in protein-dependent splicing. Once fixed, this interaction could be elaborated by further adaptive mutations in both the catalytic domain and CTD that enabled specific binding of group I introns. Our results highlight a role for non-specific RNA binding in the evolution of RNA-binding proteins.
cysteine residue at C-terminus in 8 M urea-denatured states.
Analysis of the yeast DEAD-box nucleic acid helicase Mss116p provides a structural model for how DEAD-box proteins recognize and unwind RNA duplexes. Alan Lambowitz and colleagues have solved the structure of Mss116, a yeast DEAD-box protein, bound to double-stranded RNA and a DNA–RNA hybrid. DEAD-box proteins are nucleic acid helicases that function to unwind and remodel RNAs and RNA-protein complexes. The structure shows the enzyme in a pre-unwound state, with ATP and RNA bound to different domains; it is proposed that a conformational change brings them together during unwinding. The structure also reveals how the enzyme discriminates between A-form RNA and B-form DNA. DEAD-box proteins are the largest family of nucleic acid helicases, and are crucial to RNA metabolism throughout all domains of life1,2. They contain a conserved ‘helicase core’ of two RecA-like domains (domains (D)1 and D2), which uses ATP to catalyse the unwinding of short RNA duplexes by non-processive, local strand separation3. This mode of action differs from that of translocating helicases and allows DEAD-box proteins to remodel large RNAs and RNA–protein complexes without globally disrupting RNA structure4. However, the structural basis for this distinctive mode of RNA unwinding remains unclear. Here, structural, biochemical and genetic analyses of the yeast DEAD-box protein Mss116p indicate that the helicase core domains have modular functions that enable a novel mechanism for RNA-duplex recognition and unwinding. By investigating D1 and D2 individually and together, we find that D1 acts as an ATP-binding domain and D2 functions as an RNA-duplex recognition domain. D2 contains a nucleic-acid-binding pocket that is formed by conserved DEAD-box protein sequence motifs and accommodates A-form but not B-form duplexes, providing a basis for RNA substrate specificity. Upon a conformational change in which the two core domains join to form a ‘closed state’ with an ATPase active site, conserved motifs in D1 promote the unwinding of duplex substrates bound to D2 by excluding one RNA strand and bending the other. Our results provide a comprehensive structural model for how DEAD-box proteins recognize and unwind RNA duplexes. This model explains key features of DEAD-box protein function and affords a new perspective on how the evolutionarily related cores of other RNA and DNA helicases diverged to use different mechanisms.
In the last decade, a new class of proteins has emerged that contain a topological knot in their backbone. Although these structures are rare, they nevertheless challenge our understanding of protein folding. In this review, we provide a short overview of topologically knotted proteins with an emphasis on newly discovered structures. We discuss the current knowledge in the field, including recent developments in both experimental and computational studies that have shed light on how these intricate structures fold.
The mitochondrial DEAD-box proteins Mss116p of Saccharomyces cerevisiae and CYT-19 of Neurospora crassa are ATP-dependent helicases that function as general RNA chaperones. The helicase core of each protein precedes a C-terminal extension and a basic tail, whose structural role is unclear. Here we used small-angle X-ray scattering to obtain solution structures of the full-length proteins and a series of deletion mutants. We find that the two core domains have a preferred relative orientation in the open state without substrates, and we visualize the transition to a compact closed state upon binding RNA and adenosine nucleotide. An analysis of complexes with large chimeric oligonucleotides shows that the basic tails of both proteins are attached flexibly, enabling them to bind rigid duplex DNA segments extending from the core in different directions. Our results indicate that the basic tails of DEAD-box proteins contribute to RNA-chaperone activity by binding nonspecifically to large RNA substrates and flexibly tethering the core for the unwinding of neighboring duplexes.
Topological knots are found in a considerable number of protein structures, but it is not clear how they knot and fold within the cellular environment. We investigated the behavior of knotted protein molecules as they are first synthesized by the ribosome using a cell-free translation system. We found that newly translated knotted proteins can spontaneously self-tie and do not require the assistance of molecular chaperones to fold correctly to their trefoil-knotted structures. This process is slow but efficient, and we found no evidence of misfolded species. A kinetic analysis indicates that the knotting process is rate limiting, occurs post-translationally, and is specifically and significantly (P < 0.001) accelerated by the GroEL-GroES chaperonin complex. This demonstrates a new active mechanism for this molecular chaperone and suggests that chaperonin-catalyzed knotting probably dominates in vivo. These results explain how knotted protein structures have withstood evolutionary pressures despite their topological complexity.
Structures that contain a knot formed by the path of the polypeptide backbone represent some of the most complex topologies observed in proteins. How or why these topological knots arise remains unclear. By developing a method to experimentally trap and detect knots in nonnative polypeptide chains, we find that two knotted methyltransferases, YibK and YbeA, can exist in a trefoil-knot conformation even in their chemically unfolded states. The unique denatured-state topology of these molecules explains their ability to efficiently fold to their native knotted structures in vitro and offers insights into the potential role of knots in proteins. Furthermore, the high prevalence of the denatured-state knots identified here suggests that they are either difficult to untie or that threading of any untied molecules is rapid and spontaneous. The occurrence of such knotted topologies in unfolded polypeptide chains raises the possibility that they could play an important, and as yet unexplored, role in folding and misfolding processes in vivo.
Proteins possessing deeply embedded topological knots in their structure add a stimulating new challenge to the already complex protein-folding problem. The most complicated knotted topology observed to date belongs to the human enzyme ubiquitin C-terminal hydrolase UCH-L3, which is an integral part of the ubiquitin-proteasome system. The structure of UCH-L3 contains five distinct crossings of its polypeptide chain, and it adopts a 5(2)-knotted topology, making it a fascinating target for folding studies. Here, we provide the first in depth characterization of the stability and folding of UCH-L3. We show that the protein can unfold and refold reversibly in vitro without the assistance of molecular chaperones, demonstrating that all the information necessary for the protein to find its knotted native structure is encoded in the amino acid sequence, just as with any other globular protein, and that the protein does not enter into any deep kinetic traps. Under equilibrium conditions, the unfolding of UCH-L3 appears to be two-state, however, multiphasic folding and unfolding kinetics are observed and the data are consistent with a folding pathway in which two hyperfluorescent intermediates are formed. In addition, a very slow phase in the folding kinetics is shown to be limited by proline-isomerization events. Overall, the data suggest that a knotted topology, even in its most complex form, does not necessarily limit folding in vitro, however, it does seem to require a complex folding mechanism which includes the formation of several distinct intermediate species.
Proteins possessing deeply embedded topological knots in their structure add a stimulating new challenge to the already complex protein‐folding problem. The most complicated knotted topology observed to date belongs to the human enzyme ubiquitin C‐terminal hydrolase UCH‐L3, which is an integral part of the ubiquitin–proteasome system. The structure of UCH‐L3 contains five distinct crossings of its polypeptide chain, and it adopts a 52‐knotted topology, making it a fascinating target for folding studies. Here, we provide the first in depth characterization of the stability and folding of UCH‐L3. We show that the protein can unfold and refold reversibly in vitro without the assistance of molecular chaperones, demonstrating that all the information necessary for the protein to find its knotted native structure is encoded in the amino acid sequence, just as with any other globular protein, and that the protein does not enter into any deep kinetic traps. Under equilibrium conditions, the unfolding of UCH‐L3 appears to be two‐state, however, multiphasic folding and unfolding kinetics are observed and the data are consistent with a folding pathway in which two hyperfluorescent intermediates are formed. In addition, a very slow phase in the folding kinetics is shown to be limited by proline‐isomerization events. Overall, the data suggest that a knotted topology, even in its most complex form, does not necessarily limit folding in vitro, however, it does seem to require a complex folding mechanism which includes the formation of several distinct intermediate species.