The ConSurf web-sever for the analysis of proteins, RNA, and DNA provides a quick and accurate estimate of the per-site evolutionary rate among homologues. The analysis reveals functionally important regions, such as catalytic and ligand-binding sites, which often evolve slowly. Since the last report in 2016, ConSurf has been improved in multiple ways. It now has a user-friendly interface that makes it easier to perform the analysis and to visualize the results. Evolutionary rates are calculated based on a set of homologous sequences, collected using hidden Markov model-based search tools, recently embedded in the pipeline. Using these, and following the removal of redundancy, ConSurf assembles a representative set of effective homologues for protein and nucleic acid queries to enable informative analysis of the evolutionary patterns. The analysis is particularly insightful when the evolutionary rates are mapped on the macromolecule structure. In this respect, the availability of AlphaFold model structures of essentially all UniProt proteins makes ConSurf particularly relevant to the research community. The UniProt ID of a query protein with an available AlphaFold model can now be used to start a calculation. Another important improvement is the Python re-implementation of the entire computational pipeline, making it easier to maintain. This Python pipeline is now available for download as a standalone version. We demonstrate some of ConSurf's key capabilities by the analysis of caveolin-1, the main protein of membrane invaginations called caveolae.
Microbially produced electrically conductive protein filaments are of interest because they can function as conduits for long-range biological electron transfer. They also show promise as sustainably produced electronic materials. Until now, microbially produced conductive protein filaments have been reported only for bacteria. We report here that the archaellum of Methanospirillum hungatei is electrically conductive. This is the first demonstration that electrically conductive protein filaments have evolved in Archaea Furthermore, the structure of the M. hungatei archaellum was previously determined (N. Poweleit, P. Ge, H. N. Nguyen, R. R. O. Loo, et al., Nat Microbiol 2:16222, 2016, https://doi.org/10.1038/nmicrobiol.2016.222). Thus, the archaellum of M. hungatei is the first microbially produced electrically conductive protein filament for which a structure is known. We analyzed the previously published structure and identified a core of tightly packed phenylalanines that is one likely route for electron conductance. The availability of the M. hungatei archaellum structure is expected to substantially advance mechanistic evaluation of long-range electron transport in microbially produced electrically conductive filaments and to aid in the design of "green" electronic materials that can be microbially produced with renewable feedstocks.IMPORTANCE Microbially produced electrically conductive protein filaments are a revolutionary, sustainably produced, electronic material with broad potential applications. The design of new protein nanowires based on the known M. hungatei archaellum structure could be a major advance over the current empirical design of synthetic protein nanowires from electrically conductive bacterial pili. An understanding of the diversity of outer-surface protein structures capable of electron transfer is important for developing models for microbial electrical communication with other cells and minerals in natural anaerobic environments. Extracellular electron exchange is also essential in engineered environments such as bioelectrochemical devices and anaerobic digesters converting wastes to methane. The finding that the archaellum of M. hungatei is electrically conductive suggests that some archaea might be able to make long-range electrical connections with their external environment.
Here we report that the archaellum ofMethanospirillum hungateiis electrically conductive. Our analysis of the previously published archaellum structure suggests that a core of tightly packed phenylalanines is one likely route for electron conductance. This is the first demonstration that electrically conductive protein filaments (e-PFs) have evolved in Archaea and is the first e-PF for which a structure is known, facilitating mechanistic evaluation of long-range electron transport in e-PFs.
The metallic-like electrical conductivity of Geobacter sulfurreducens pili has been documented with multiple lines of experimental evidence, but there is only a rudimentary understanding of the structural features which contribute to this novel mode of biological electron transport. In order to determine if it was feasible for the pilin monomers of G. sulfurreducens to assemble into a conductive filament, theoretical energy-minimized models of Geobacter pili were constructed with a previously described approach, in which pilin monomers are assembled using randomized structural parameters and distance constraints. The lowest energy models from a specific group of predicted structures lacked a central channel, in contrast to previously existing pili models. In half of the no-channel models the three N-terminal aromatic residues of the pilin monomer are arranged in a potentially electrically conductive geometry, sufficiently close to account for the experimentally observed metallic like conductivity of the pili that has been attributed to overlapping pi-pi orbitals of aromatic amino acids. These atomic resolution models capable of explaining the observed conductive properties of Geobacter pili are a valuable tool to guide further investigation of the metallic-like conductivity of the pili, their role in biogeochemical cycling, and applications in bioenergy and bioelectronics.
The degree of evolutionary conservation of an amino acid in a protein or a nucleic acid in DNA/RNA reflects a balance between its natural tendency to mutate and the overall need to retain the structural integrity and function of the macromolecule. The ConSurf web server (http://consurf.tau.ac.il), established over 15 years ago, analyses the evolutionary pattern of the amino/nucleic acids of the macromolecule to reveal regions that are important for structure and/or function. Starting from a query sequence or structure, the server automatically collects homologues, infers their multiple sequence alignment and reconstructs a phylogenetic tree that reflects their evolutionary relations. These data are then used, within a probabilistic framework, to estimate the evolutionary rates of each sequence position. Here we introduce several new features into ConSurf, including automatic selection of the best evolutionary model used to infer the rates, the ability to homology-model query proteins, prediction of the secondary structure of query RNA molecules from sequence, the ability to view the biological assembly of a query (in addition to the single chain), mapping of the conservation grades onto 2D RNA models and an advanced view of the phylogenetic tree that enables interactively rerunning ConSurf with the taxa of a sub-tree.
ABSTRACT Direct measurement of multiple physical properties of Geobacter sulfurreducens pili have demonstrated that they possess metallic-like conductivity, but several studies have suggested that metallic-like conductivity is unlikely based on the structures of the G. sulfurreducens pilus predicted from homology models. In order to further evaluate this discrepancy, pili were examined with synchrotron X-ray microdiffraction and rocking-curve X-ray diffraction. Both techniques revealed a periodic 3.2-Å spacing in conductive, wild-type G. sulfurreducens pili that was missing in the nonconductive pili of strain Aro5, which lack key aromatic acids required for conductivity. The intensity of the 3.2-Å peak increased 100-fold when the pH was shifted from 10.5 to 2, corresponding with a previously reported 100-fold increase in pilus conductivity with this pH change. These results suggest a clear structure-function correlation for metallic-like conductivity that can be attributed to overlapping π-orbitals of aromatic amino acids. A homology model of the G. sulfurreducens pilus was constructed with a Pseudomonas aeruginosa pilus model as a template as an alternative to previous models, which were based on a Neisseria gonorrhoeae pilus structure. This alternative model predicted that aromatic amino acids in G. sulfurreducens pili are packed within 3 to 4 Å, consistent with the experimental results. Thus, the predictions of homology modeling are highly sensitive to assumptions inherent in the model construction. The experimental results reported here further support the concept that the pili of G. sulfurreducens represent a novel class of electronically functional proteins in which aromatic amino acids promote long-distance electron transport. IMPORTANCE The mechanism for long-range electron transport along the conductive pili of Geobacter sulfurreducens is of interest because these “microbial nanowires” are important in biogeochemical cycling as well as applications in bioenergy and bioelectronics. Although proteins are typically insulators, G. sulfurreducens pilus proteins possess metallic-like conductivity. The studies reported here provide important structural insights into the mechanism of the metallic-like conductivity of G. sulfurreducens pili. This information is expected to be useful in the design of novel bioelectronic materials.
easier to use than either Chime or RasMol. Moreover, it is much more powerful; in addition to basic macromolecular visualization capabilities common to most similar programs, it offers one-click visualization of interfaces between moieties ('contacts'), cation–π π interactions and salt bridges, as well as easy-to-use routines to visualize regions of conservation in three-dimensional protein structures based on multiple sequence alignments. Protein Explorer (PE) is built upon Chime, a molecular graphics browser plugin that is freeware from MDL Information Systems (www.mdlchime.com). It was possible to implement PE within a few years only because of the power inherent in Chime. Chime, in turn, is in part built upon the molecular graphics rendering and command language in RasMol [1,2]. However, Chime has several additional significant capabilities, such as the ability to render solvent-accessible molecular surfaces and animations. The problem is that to get much out of either Chime or RasMol, the user must learn a complicated and extensive command language. This requirement makes the power of Chime and RasMol inaccessible to most of those who could benefit from macromolecular visualization. PE addresses this problem by enabling both basic and complex visualizations from menus, buttons and forms, without requiring the user to learn a single RasMol-style command. Nevertheless, PE accepts RasMol commands as a convenience for those who have learned them (see command input slot in Fig. 1). A detailed ease-of-use comparison of PE with RasMol is available at on-line, or downloaded for off-line use. Web links can pre-specify the molecules to be displayed, supporting course or textbook websites. PE runs in the Netscape browser (Windows or Macintosh), or in Microsoft Internet Explorer (Windows only). This article is not meant to provide instructions for the use of PE because the program has extensive built-in instructions. Rather, it is intended to help readers decide whether PE will be useful in their work. The first image of any molecule shown by PE is designed to be maximally informative. It is accompanied by a generic description, with links to illustrated explanations of backbone traces, disulfide bonds, 'hetero atoms', a standard color scheme for identifying elements, the absence or presence of hydrogen atoms, and water in protein crystals. (Typically only 10–20% of the water is tightly enough bound to be resolved and displayed. This kind of information is available through links on the FirstView page.) Clicking on any atom reports its element, and the name and sequence number …
Proteopedia ( http://proteopedia.org ) is an interactive web resource with 3D rotating models that react and change following user interaction. The main goals of Proteopedia are to collect, organize and disseminate structural and functional knowledge about protein, RNA, DNA, and other macromolecules, and their assemblies and interactions with small molecules, in a manner that is relevant and broadly accessible to students and scientists. This Guide provides instructions for the potential author of pages in Proteopedia, and for educators and teachers in incorporating Proteopedia when teaching proteins, and protein-functions. Of course you can use Proteopedia as a reference resource without authoring content.
Many mutations disappear from the population because they impair protein function and/or stability. Thus, amino acid positions that are essential for proper function evolve more slowly than others, or in other words, the slow evolutionary rate of a position reflects its importance. Con- Surf (http://consurf.tau.ac.il), reviewed in this manuscript, exploits this to reveal key amino acid positions that are im- portant for maintaining the native conformation(s) of the protein and its function, be it binding, catalysis, transport, etc. Given the sequence or 3D structure of the query protein as input, a search for similar sequences is conducted and the sequences are aligned. The multiple sequence alignment is subsequently used to calculate the evolutionary rates of each amino acid site, using Bayesian or maximum-likelihood algorithms. Both algorithms take into account the evolution- ary relationships between the sequences, reflected in phylo- genetic trees, to alleviate problems due to uneven (biased) sampling in sequence space. This is particularly important when the number of sequences is low. The ConSurf-DB, a new release of which is presented here, provides precalcu- lated ConSurf conservation analysis of nearly all available structures in the Protein DataBank (PDB). The usefulness of ConSurf for the study of individual proteins and mutations, as well as a range of large-scale, genome-wide applications, is reviewed.
The cover shows on the left a ConSurf analysis of the influenza neuraminidase protein. The 3D tetramer colored by conservation grades is shown in (A) with maroon indicating the most conserved amino acids, and close-up views (B,C) are shown of one of the conserved regions with the anti-flu drug, oseltamivir (tamiflu) bound; see Celniker et al. for details, page 199. On the right is a page from Proteopedia as viewed on an iPad, via the JSmol viewer, showing a complex of HIV-Protease and Saquinavir (Invirase), the first protease inhibitor approved by the FDA for the treatment of HIV; see Hanson et al. for details, page 207
3D visualization assists in identifying diverse mechanisms of protein-DNA recognition that can be observed for transcription factors and other DNA binding proteins. We used Proteopedia to illustrate transcription factor-DNA readout modes with a focus on DNA shape, which can be a function of either nucleotide sequence (Hox proteins) or base pairing geometry (p53). © 2012 by The International Union of Biochemistry and Molecular Biology.
, Amit Kessel and , Nir Ben-Tal, CRC Press, Boca Raton, FL, USA, 2011, 626 pp., ISBN 978-1-4398-1071-2 (hardcover, $79.95). Also available as e-book. Eric Martz*, * Department of Microbiology, University of Massachusetts, Amherst, Massachusetts 01003. Introduction to Proteins by Kessel and Ben-Tal is an excellent, state-of-the-art choice for students, faculty, or researchers needing a monograph on protein structure. It is unusual, among monographs dealing primarily with structure versus function, in including protein dynamics and energetics in considerable detail. It is also unusual in its strong, detailed coverage of membrane proteins (81 pages, Chapter 7), 20 pages on post-translational modifications, and 14 pages on intrinsically unstructured proteins. G-protein coupled receptors are given thorough coverage (23 pages). The book is clear, well organized, aptly illustrated in color, and a pleasure to read. The first two chapters are an impressive textbook unto themselves: 190 pages introducing “the importance of proteins in living organisms,” including a useful review of the biochemistry of living systems, diversity of proteins and their functions. The book continues by reviewing non-covalent bonds, pKa, primary, secondary, tertiary and quaternary structure, and post-translational modifications. Particularly excellent are the introductions to amyloidoses and protein misfolding, chaperonins, allostery, and hemoglobin. Globular proteins are covered amply throughout the book, and there are 25 pages on fibrous proteins emphasizing function-structure relationships, with examples primarily from eukaryotic cells. The final chapter deals with protein-ligand (including protein-protein) interactions, including a clear discussion of lock-and-key, induced fit, and population shift models, as well as free energies of binding. Chapter 3 offers a 43 page overview of “methods of structure determination and prediction.” Although useful as far as it goes, this is one of the weaker chapters. It does not emphasize that the protein must be soluble for all common methods of structure determination, does not mention the use of heavy metals or selenomethionine for phase determination, and fails to offer much guidance on evaluation of the quality of crystallographic or NMR models. Free R and temperature factor (B factor) are not mentioned. Thought-provoking opinions are occasionally stated. For example, after introducing SCOP and CATH: “The problem of finding a single foolproof classification method may result … from the protein structural space being continuous …. even apparently different structures may share common features such as secondary structure arrangement. If this is indeed true, classifying proteins into discrete structural categories may be pointless.” (p. 157–8) The book is thoroughly documented with citations to the literature gathered at the end of each chapter. All references include titles. For example, Chapter 7 on membrane proteins cites 400 references, including many as recent as 2009. The entire book cites nearly 2,000 references. PDB identification codes are frequently given for the examples described, making it easy to visualize them in 3D. Unfortunately, no guidance is given regarding visualization software. The book's website offers only a set of study questions, about a dozen for each of the eight chapters. These are generally thought-provoking and useful, and answers, as well as slides, are said to be available to instructors on request from the publisher (p. xxiii). There are a few topics that deserved greater coverage. Not conveyed are the prevalence of cation-pi interactions, and their importance in the energetics of protein folding and in ligand-protein interactions (e.g., acetylcholinesterase, nicotine addiction). The absence of cation-pi interactions in the introduction to noncovalent interactions is conspicuous. An otherwise extensive discussion of hemoglobin structure versus function omits mention of malaria as an evolutionary pressure maintaining sickle hemoglobin. Although protein engineering is introduced briefly, some coverage of recent successes with enzyme re-targeting would be useful. There are rare errors, such as describing lipids as macromolecules, and lumping the effect of mercaptoethanol in Anfinson's denaturation of RNAse under “disrupting noncovalent forces” (p. 103). The English is excellent, with very rare minor slips, such as calling the ends of protein chains their “edges” or the thickness of a lipid bilayer its “width.” Although extremophiles are introduced briefly (p. 288), I did not see mention of the frequent successes crystallographers have had with proteins from thermophiles. Although mechanisms of thermostabilization are discussed, I saw no mention of the role of salt bridges. The terms thermophile, thermostability, and hyperthermophile did not make it into the index, although the index is 34 pages long and generally quite useful. Many of the many acronyms used throughout the book made it into the index, but some did not (e.g., PFAM, CL). The book concludes with an extensive treatment of ligand-protein interactions, including 30 pages on drug action and design. A clear 20-page overview and explanation of various techniques is followed by a 10-page case study of the rational design of angiotensin conversion enzyme inhibitors. Overall, this is an immensely informative, thoroughly researched, up to date text, with broad coverage and remarkable depth. Introduction to Proteins would provide an excellent basis for an upper level or graduate course on protein structure, and a valuable addition to the libraries of professionals interested in this centrally important field.
Proteopedia is a collaborative, 3D web-encyclopedia of protein, nucleic acid and other biomolecule structures. Created as a means for communicating biomolecule structures to a diverse scientific audience, Proteopedia (http://www.proteopedia.org) presents structural annotation in an intuitive, interactive format and allows members of the scientific community to easily contribute their own annotations. Here, we provide a status report on Proteopedia by describing advances in the web resource since its inception three and a half years ago, focusing on features of potential direct use to the scientific community. We discuss its progress as a collaborative 3D-encyclopedia of structures as well as its use as a complement to scientific publications and PowerPoint presentations. We also describe Proteopedia's use for 3D visualization in structure-related pedagogy.
It is informative to detect highly conserved positions in proteins and nucleic acid sequence/structure since they are often indicative of structural and/or functional importance. ConSurf (http://consurf.tau.ac.il) and ConSeq (http://conseq.tau.ac.il) are two well-established web servers for calculating the evolutionary conservation of amino acid positions in proteins using an empirical Bayesian inference, starting from protein structure and sequence, respectively. Here, we present the new version of the ConSurf web server that combines the two independent servers, providing an easier and more intuitive step-by-step interface, while offering the user more flexibility during the process. In addition, the new version of ConSurf calculates the evolutionary rates for nucleic acid sequences. The new version is freely available at: http://consurf.tau.ac.il/.
Richard C. Garratt and Christine A. Orengo, Wiley-VCH, 2008, ISBN: 978-3-527-31963-3, $19.99 or €14.90. A six-page plastic-laminated reference chart in color on lightweight A4 cardstock (threefold-out double-sided panels, 11 × 8.3 inches/28 × 21 cm). Eric Martz*, * Department of Microbiology, University of Massachusetts, Amherst, Massachusetts 01003. Suitable for college biochemistry students as well as biochemical educators and researchers, The Protein Chart packs an astonishing amount of information about protein 3D tertiary and quaternary structure into a well-organized, compact six-page reference chart. The title is misleading: The Protein Chart covers only a relatively narrow aspect of protein science. Derived from the CATH protein structure classification, the Chart fills five pages, with an attractively designed cover occupying the sixth. “CATH is a hierarchical classification of protein domain structures, which clusters proteins at four major levels, Class (C), Architecture (A), Topology (T), and Homologous superfamily (H)” (www.cathdb.info). The bulk of the chart, three pages, is devoted to tertiary domain structure. There are 26 domain architectures (columns), each illustrated with 1–5 structures representing fold groups (rows). Each example shows a ribbon schematic, colored by secondary structure. The domain architectures are divided into four classes: alpha proteins (17 examples), beta proteins (28 examples), alpha/beta proteins (32 examples), and knots and fibers (nine examples). These three pages thus contain 86 examples of fold groups in a hierarchical tabular format. Each example includes the fold name, the PDB (Protein Data Bank) identification code, chain identifier within the PDB entry and domain number, the CATH code for fold group, a “secondary structure string”, an average length in residues with standard deviation, and an indication of whether the domain occurs in eukarya, archaea, or bacteria. The functions of the molecules containing the members of each fold group are also indicated, using eight categories of functions. A full page is devoted to quaternary structure (oligomers), divided into columns according to rotational symmetry – the number of identical subunits per 360 degree turn. Four to five examples are in each column, annotated with the name of the protein, the PDB identifier, the number of subunits, and the point group symmetry (in both Schoenflies and international nomenclature), and the presence or absence of dihedral symmetry. The contributions of oligomerization to structure or function are also indicated, using eight categories. A half page is devoted to 10 examples of basic topologies of secondary structure (Greek key, jelly roll, immunoglobulin domain, Rossman fold, etc.), and a half page to important structural motifs (helix-turn-helix, EF-hand, leucine zipper, zinc finger, etc.). There are some minor inadequacies, which perhaps could be addressed in a second edition. A few things were unclear, and I wished that some of the blank areas (nearly one-quarter of the three pages on domains is blank) had been used to explain and clarify. Alternatively, a website with explanatory notes would be very welcome. (No website was mentioned except cathdb.info, where I found no mention of The Protein Chart.) From the following statement, I could not tell what the tabulated percentages represent: “The population given as a percentage for each architecture is calculated from the 527 genomes present in Gene3D version 6.0.” The function “binding” seemed to deserve some clarification, as did “domain number”. Is the “number of subunits” (on the oligomer page) synonymous with “number of protein chains”? It would have been interesting to know how many of the 37 oligomer examples were homo-oligomers (most, it appeared to me, with one hetero-oligomer example being the proteasome 1pma). Also, as average lengths were given for each of the 86-fold group examples, it would have been interesting to know the total number of structures in each fold group. The uninitiated could be left with the impression that this chart gives an overview of all that is known about protein 3D structure. A more complete impression could have been achieved had the widespread occurrence of intrinsically unstructured protein been mentioned, illustrated with a few examples of sequences that adopt stable uniform structures only when folding against a previously folded domain. The Protein Chart offers an overview and sampling of the breadth of protein 3D tertiary and quaternary structural knowledge and classification that is eminently useful and at the same time quite beautiful. The convenient inclusion of PDB accession codes makes it easy to further explore any of the 123 examples. There is plenty to absorb and to ponder both for beginners and for structural biologists. It is difficult to imagine any better way to convey the state of the art so graphically and succinctly.
BACKGROUND:Detecting candidate B-cell epitopes in a protein is a basic and fundamental step in many immunological applications. Due to the impracticality of experimental approaches to systematically scan the entire protein, a computational tool that predicts the most probable epitope regions is desirable.RESULTS:The Epitopia server is a web-based tool that aims to predict immunogenic regions in either a protein three-dimensional structure or a linear sequence. Epitopia implements a machine-learning algorithm that was trained to discern antigenic features within a given protein. The Epitopia algorithm has been compared to other available epitope prediction tools and was found to have higher predictive power. A special emphasis was put on the development of a user-friendly graphical interface for displaying the results.CONCLUSION:Epitopia is a user-friendly web-server that predicts immunogenic regions for both a protein structure and a protein sequence. Its accuracy and functionality make it a highly useful tool. Epitopia is available at http://epitopia.tau.ac.il and includes extensive explanations and example predictions.
Page s 39 information [1].Proteopedia displays protein structures and other biomacromolecules interactively.These 3D interactive images can be rotated and zoomed, and are surrounded by text with hyperlinks that change the appearance of the 3D structure to reflect the concept explained in the text.This makes the complex structural information readily accessible and comprehensible, even to non-structural biologists.Using Proteopedia, anyone can easily create descriptions of biomacromolecules linked to their 3D structures, e.g.: (a) Proton Channels: http://proteopedia.org/wiki/index.php/Proton_Channels(b) HIV-1 protease: http://proteopedia.org/wiki/index.php/HIV-1_protease(c) Beta-adrenergic receptor: http://www.proteopedia.org/wiki/index.php/
Many scientists lack the background to fully utilize the wealth of solved three-dimensional biomacromolecule structures. Thus, a resource is needed to present structure/function information in a user-friendly manner to a broad scientific audience. Proteopedia http://www.proteopedia.org is an interactive, wiki web-resource whose pages have embedded three-dimensional structures surrounded by descriptive text containing hyperlinks that change the appearance (view, representations, colors, labels) of the adjacent three-dimensional structure to reflect the concept explained in the text.
Roded Sharan合作论文数Tel-Aviv University;School of Computer Science2