The heterologous expression of integral membrane proteins (IMPs) remains a major bottleneck in the characterization of this important protein class. IMP expression levels are currently unpredictable, which renders the pursuit of IMPs for structural and biophysical characterization challenging and inefficient. Experimental evidence demonstrates that changes within the nucleotide or amino acid sequence for a given IMP can dramatically affect expression levels, yet these observations have not resulted in generalizable approaches to improve expression levels. Here, we develop a data-driven statistical predictor named IMProve that, using only sequence information, increases the likelihood of selecting an IMP that expresses in Escherichia coli. The IMProve model, trained on experimental data, combines a set of sequence-derived features resulting in an IMProve score, where higher values have a higher probability of success. The model is rigorously validated against a variety of independent data sets that contain a wide range of experimental outcomes from various IMP expression trials. The results demonstrate that use of the model can more than double the number of successfully expressed targets at any experimental scale. IMProve can immediately be used to identify favorable targets for characterization. Most notably, IMProve demonstrates for the first time that IMP expression levels can be predicted directly from sequence.
Diabetes research studies routinely rely upon the use of tissue samples from human organ donors. It remains unclear whether the length of hospital stay prior to organ donation affects the presence of cells infiltrating the pancreas or the frequency of replicating beta cells.
The expression of integral membrane proteins (IMPs) remains a major bottleneck in the characterization of this important protein class. IMP expression levels are currently unpredictable, which renders the pursuit of IMPs for structural and biophysical characterization challenging and inefficient. Experimental evidence demonstrates that changes within the nucleotide or amino-acid sequence for a given IMP can dramatically affect expression; yet these observations have not resulted in generalizable approaches to improved expression. Here, we develop a data-driven statistical predictor named IMProve, that, using only sequence information, increases the likelihood of selecting an IMP that expresses in E. coli. The IMProve model, trained on experimental data, combines a set of sequence-derived features resulting in an IMProve score, where higher values have a higher probability of success. The model is rigorously validated against a variety of independent datasets that contain a wide range of experimental outcomes from various IMP expression trials. The results demonstrate that use of the model can more than double the number of successfully expressed targets at any experimental scale. IMProve can immediately be used to identify favorable targets for characterization.
Membrane protein production is difficult; their biogenesis does not stop with translation but also requires translocation and integration into a lipid bilayer. These additional steps hamper their heterologous expression which significantly impedes biophysical and structural studies. Detailed and anecdotal evidence in literature suggests that a variety nucleotide and amino-acid sequence level determinants may potentially support or hinder their biogenesis, e.g. mRNA pausing elements, codon adaptation, transmembrane segment hydrophobicity, “positive inside rule.” By training a preference-ranking Support Vector Machine, we have developed a statistical model that predicts a relative likelihood of a membrane protein's successful expression using quantitative experimental data of overexpression. This model is rigorously validated against expression outcomes from small-scale laboratory experiments (e.g. expression tests that routinely precede structural studies) published in the literature as well as large-scale expression trials from a Protein Structure Initiative consortium facility. We show remarkable agreement between the predicted and experimental expression outcomes and propose our model, trained and cross-validated on the entire corpus of data, as a tool for the membrane protein biophysics community to streamline the process of overexpressing a target for study. Given the framework of our model, it can be trivially re-trained as additional experimental outcomes are gathered from past work or created from experiments. Furthermore, the relative weights gathered from parameters of the statistical model may help further characterize translocation mechanisms and suggest intriguing areas for further biophysical and computational experiments.
Integral membrane proteins (IMPs) control the flow of information and nutrients across cell membranes, yet IMP mechanistic studies are hindered by difficulties in expression. We investigate this issue by addressing the connection between IMP sequence and observed expression levels. For homologs of the IMP TatC, observed expression levels vary widely and are affected by small changes in protein sequence. The effect of sequence changes on experimentally observed expression levels strongly correlates with the simulated integration efficiency obtained from coarse-grained modeling, which is directly confirmed using an in vivo assay. Furthermore, mutations that improve the simulated integration efficiency likewise increase the experimentally observed expression levels. Demonstration of these trends in both Escherichia coli and Mycobacterium smegmatis suggests that the results are general to other expression systems. This work suggests that IMP integration is a determinant for successful expression, raising the possibility of controlling IMP expression via rational design.
Mechanosensitive channels are ubiquitous and highly studied. However, the evolution of the bacterial channels remains enigmatic. It can be argued that mechanosensitivity might be a feature of all membrane proteins with some becoming progressively less sensitive to membrane tension over the course of evolution. Bacteria and archaea exhibit two main classes of channels, MscS and MscL. Present day channels suggest that the evolution of MscL may be highly constrained, whereas MscS has undergone elaboration via gene fusion (and potentially gene fission) events to generate a diversity of channel structures. Some of these channel variants are constrained to a small number of genera or species. Some are only found in higher organisms. Only exceptionally have these diverse channels been investigated in any detail. In this review we consider both the processes that might have led to the evolved complexity but also some of the methods exploiting the explosion of genome sequences to understand (and/or track) their distribution. The role of MscS-related channels in calcium-mediated cell biology events is considered.
In Campylobacterales and related ε-proteobacteria with N-linked glycosylation (NLG) pathways, free oligosaccharides (fOS) are released into the periplasmic space from lipid-linked precursors by the bacterial oligosaccharyltransferase (PglB). This hydrolysis results in the same molecular structure as the oligosaccharide that is transferred to a protein to be glycosylated. This allowed for the general elucidation of the fOS-branched structures and monosaccharides from a number of species using standard enrichment and mass spectrometry methods. To aid characterization of fOS, hydrazide chemistry has often been used for chemical modification of the reducing part of oligosaccharides resulting in better selectivity and sensitivity in mass spectrometry; however, the removal of the unreacted reagents used for the modification often causes the loss of the sample. Here, we develop a more robust method for fOS purification and characterize glycostructures using complementary tandem mass spectrometry (MS/MS) analysis. A cationic cysteine hydrazide derivative was synthesized to selectively isolate fOS from periplasmic fractions of bacteria. The cysteine hydrazide nicotinamide (Cyhn) probe possesses both thiol and cationic moieties. The former enables reversible conjugation to a thiol-activated solid support, while the latter improves the ionization signal during MS analysis. This enrichment was validated on the well-studied Campylobacter jejuni by identifying fOS from the periplasmic extracts. Using complementary MS/MS analysis, we approximated data of a known structure of the fOS from Campylobacter concisus. This versatile enrichment technique allows for the exploration of a diversity of protein glycosylation pathways.
ATP-binding cassette (ABC) transporters, although being ubiquitous in biology, often feature a subunit that is limited primarily to bacteria and archaea. This subunit, the substrate-binding protein (SBP), is a key determinant of the substrate specificity and high affinity of ABC uptake systems in these organisms. Most prokaryotes have many SBP-dependent ABC transporters that recognize a broad range of ligands from metal ions to amino acids, sugars and peptides. Herein, we review the structure and function of a number of more unusual SBPs, including an ABC transporter involved in the transport of rare furanose forms of sugars and an SBP that has evolved to specifically recognize the bacterial cell wall-derived murein tripeptide (Mtp). Both these examples illustrate that subtle changes in binding-site architecture, including changes in side chains not directly involved in ligand co-ordination, can result in significant alteration of substrate range in novel and unpredictable ways.
Campylobacter jejuni is one of the most successful food‐borne human pathogens. Here we use electron cryotomography to explore the ultrastructure of C. jejuni cells in logarithmically growing cultures. This provides the first look at this pathogen in a near‐native state at macromolecular resolution (~5 nm). We find a surprisingly complex polar architecture that includes ribosome exclusion zones, polyphosphate storage granules, extensive collar‐shaped chemoreceptor arrays, and elaborate flagellar motors.
The bacterial flagellum is one of nature's most amazing and well‐studied nanomachines. Its cell‐wall‐anchored motor uses chemical energy to rotate a microns‐long filament and propel the bacterium towards nutrients and away from toxins. While much is known about flagellar motors from certain model organisms, their diversity across the bacterial kingdom is less well characterized, allowing the occasional misrepresentation of the motor as an invariant, ideal machine. Here, we present an electron cryotomographical survey of flagellar motor architectures throughout the Bacteria. While a conserved structural core was observed in all 11 bacteria imaged, surprisingly novel and divergent structures as well as different symmetries were observed surrounding the core. Correlating the motor structures with the presence and absence of particular motor genes in each organism suggested the locations of five proteins involved in the export apparatus including FliI, whose position below the C‐ring was confirmed by imaging a deletion strain. The combination of conserved and specially‐adapted structures seen here sheds light on how this complex protein nanomachine has evolved to meet the needs of different species. A comprehensive electron cryotomographical survey of bacterial flagellar motors reveals the existence of a conserved structural core that is surrounded by a divergent set of novel structural features. Key proteins of the flagellar export apparatus can now be localized within the motor.
Chemoreceptors are key components of the high-performance signal transduction system that controls bacterial chemotaxis. Chemoreceptors are typically localized in a cluster at the cell pole, where interactions among the receptors in the cluster are thought to contribute to the high sensitivity, wide dynamic range, and precise adaptation of the signaling system. Previous structural and genomic studies have produced conflicting models, however, for the arrangement of the chemoreceptors in the clusters. Using whole-cell electron cryo-tomography, here we show that chemoreceptors of different classes and in many different species representing several major bacterial phyla are all arranged into a highly conserved, 12-nm hexagonal array consistent with the proposed “trimer of dimers” organization. The various observed lengths of the receptors confirm current models for the methylation, flexible bundle, signaling, and linker sub-domains in vivo. Our results suggest that the basic mechanism and function of receptor clustering is universal among bacterial species and was thus conserved during evolution.
The widespread utilization of sugars by microbes is reflected in the diversity and multiplicity of cellular transporters used to acquire these compounds from the environment. The model bacterium Escherichia coli has numerous transporters that allow it to take up hexoses and pentoses, which recognize the more abundant pyranose forms of these sugars. Here we report the biochemical and structural characterization of a transporter protein YtfQ from E. coli that forms part of an uncharacterized ABC transporter system. Remarkably the crystal structure of this protein, solved to 1.2 A using x-ray crystallography, revealed that YtfQ binds a single molecule of galactofuranose in its ligand binding pocket. Selective binding of galactofuranose over galactopyranose was also observed using NMR methods that determined the form of the sugar released from the protein. The pattern of expression of the ytfQRTyjfF operon encoding this transporter mirrors that of the high affinity galactopyranose transporter of E. coli, suggesting that this bacterium has evolved complementary transporters that enable it to use all the available galactose present during carbon limiting conditions.
The acquisition of host-derived sialic acid is an important virulence factor for some bacterial pathogens, but in vivo this sugar acid is sequestered in sialoconjugates as the alpha-anomer. In solution, however, sialic acid is present mainly as the beta-anomer, formed by a slow spontaneous mutarotation. We studied the Escherichia coli protein YjhT as a member of a family of uncharacterized proteins present in many sialic acid-utilizing pathogens. This protein is able to accelerate the equilibration of the alpha- and beta-anomers of the sialic acid N-acetylneuraminic acid, thus describing a novel sialic acid mutarotase activity. The structure of this periplasmic protein, solved to 1.5 A resolution, reveals a dimeric 6-bladed unclosed beta-propeller, the first of a bacterial Kelch domain protein. Mutagenesis of conserved residues in YjhT demonstrated an important role for Glu-209 and Arg-215 in mutarotase activity. We also present data suggesting that the ability to utilize alpha-N-acetylneuraminic acid released from complex sialoconjugates in vivo provides a physiological advantage to bacteria containing YjhT.
The PEB1a protein is an antigenic factor exposed on the surface of the food-borne human pathogen Campylobacter jejuni, which has a major role in adherence and host colonisation. PEB1a is also the periplasmic binding protein component of an aspartate/glutamate ABC transporter essential for optimal microaerobic growth on these dicarboxylic amino acids. Here, we report the crystal structure of PEB1a at 1.5 A resolution. The protein has a typical two-domain alpha/beta structure, characteristic of periplasmic extracytoplasmic solute receptors and a chain topology related to the type II subfamily. An aspartate ligand, clearly defined by electron density in the interdomain cleft, forms extensive polar interactions with the protein, the majority of which are made with the larger domain. Arg89 and Asp174 form ion-pairing interactions with the main chain alpha-carboxyl and alpha-amino-groups, respectively, of the ligand, while Arg67, Thr82, Lys19 and Tyr156 co-ordinate the ligand side-chain carboxyl group. Lys19 and Arg67 line a positively charged groove, which favours binding of Asp over the neutral Asn. The ligand-binding cleft is of sufficient depth to accommodate a glutamate. This is the first structure of an ABC-type aspartate-binding protein, and explains the high affinity of the protein for aspartate and glutamate, and its much weaker binding of asparagine and glutamine. Stopped-flow fluorescence spectroscopy indicates a simple bimolecular mechanism of ligand binding, with high association rate constants. Sequence alignments and phylogenetic analyses revealed PEB1a homologues in some Gram-positive bacteria. The alignments suggest a more distant homology with GltI from Escherichia coli, a known glutamate and aspartate-binding protein, but Lys19 and Tyr156 are not conserved in GltI. Our results provide a structural basis for understanding both the solute transport and adhesin/virulence functions of PEB1a.
Mittendrin: Die zweikernige Eisen(III)-Enterobactin-Modellverbindung [{Fe(mecam)}2]6− (H6-mecam=1,3,5-N,N′,N′′-Tris(2,3-dihydroxybenzoyl)triaminomethylbenzol)wird von zwei Molekülen eines periplasmatischen Bindungsproteins (gelb und blau; siehe Struktur) erkannt. Das Aggregat wird durch hydrophobe Wechselwirkungen zwischen den Liganden des zweikernigen Eisenkomplexes stabilisiert.
The Structural Proteomics In Europe ( SPINE) consortium contained a workpackage to address the automated X- ray analysis of macromolecules. The aim of this workpackage was to increase the throughput of three- dimensional structures while maintaining the high quality of conventional analyses. SPINE was able to bring together developers of software with users from the partner laboratories. Here, the results of a workshop organized by the consortium to evaluate software developed in the member laboratories against a set of bacterial targets are described. The major emphasis was on molecular- replacement suites, where automation was most advanced. Data processing and analysis, use of experimental phases and model construction were also addressed, albeit at a lower level.
A collaborative project between two Structural Proteomics In Europe ( SPINE) partner laboratories, York and Oxford, aimed at high- throughput ( HTP) structure determination of proteins from Bacillus anthracis, the aetiological agent of anthrax and a biomedically important target, is described. Based upon a target- selection strategy combining ` lowhanging fruit' and more challenging targets, this work has contributed to the body of knowledge of B. anthracis, established and developed HTP cloning and expression technologies and tested HTP pipelines. Both centres developed ligation- independent cloning ( LIC) and expression systems, employing custom LIC- PCR, Gateway and In- Fusion technologies, used in combination with parallel protein purification and robotic nanolitre crystallization screening. Overall, 42 structures have been solved by X- ray crystallography, plus two by NMR through collaboration between York and the SPINE partner in Utrecht. Three biologically important protein structures, BA4899, BA1655 and BA3998, involved in tRNA modification, sporulation control and carbohydrate metabolism, respectively, are highlighted. Target analysis by biophysical clustering based on pI and hydropathy has provided useful information for future target- selection strategies. The technological developments and lessons learned from this project are discussed. The success rate of protein expression and structure solution is at least in keeping with that achieved in structural genomics programs.