Background The Critical Assessment of Functional Annotation (CAFA) is an ongoing, global, community-driven effort to evaluate and improve the computational annotation of protein function. Results Here, we report on the results of the third CAFA challenge, CAFA3, that featured an expanded analysis over the previous CAFA rounds, both in terms of volume of data analyzed and the types of analysis performed. In a novel and major new development, computational predictions and assessment goals drove some of the experimental assays, resulting in new functional annotations for more than 1000 genes. Specifically, we performed experimental whole-genome mutation screening in Candida albicans and Pseudomonas aureginosa genomes, which provided us with genome-wide experimental data for genes associated with biofilm formation and motility. We further performed targeted assays on selected genes in Drosophila melanogaster , which we suspected of being involved in long-term memory. Conclusion We conclude that while predictions of the molecular function and biological process annotations have slightly improved over time, those of the cellular component have not. Term-centric prediction of experimental annotations remains equally challenging; although the performance of the top methods is significantly better than the expectations set by baseline methods in C. albicans and D. melanogaster , it leaves considerable room and need for improvement. Finally, we report that the CAFA community now involves a broad range of participants with expertise in bioinformatics, biological experimentation, biocuration, and bio-ontologies, working together to improve functional annotation, computational function prediction, and our ability to manage big data in the era of large experimental screens.
Background: Plasmodium falciparum malaria is one of the most widespread parasitic infections in humans and remains a leading global health concern. Malaria elimination efforts are threatened by the emergence and spread of resistance to artemisinin-based combination therapy, the first-line treatment of malaria. Promising molecular markers and pathways associated with artemisinin drug resistance have been identified, but the underlying molecular mechanisms of resistance remains unknown. The genomic data from early period of emergence of artemisinin resistance (2008-2011) was evaluated, with aim to define k13 associated genetic background in Cambodia, the country identified as epicentre of anti-malarial drug resistance, through characterization of 167 parasite isolates using a panel of 21,257 SNPs. Results: Eight subpopulations were identified suggesting a process of acquisition of artemisinin resistance consistent with an emergence-selection-diffusion model, supported by the shifting balance theory. Identification of population specific mutations facilitated the characterization of a core set of 57 background genes associated with artemisinin resistance and associated pathways. The analysis indicates that the background of artemisinin resistance was not acquired after drug pressure, rather is the result of fixation followed by selection on the daughter subpopulations derived from the ancestral population. Conclusions: Functional analysis of artemisinin resistance subpopulations illustrates the strong interplay between ubiquitination and cell division or differentiation in artemisinin resistant parasites. The relationship of these pathways with the P. falciparum resistant subpopulation and presence of drug resistance markers in addition to k13, highlights the major role of admixed parasite population in the diffusion of artemisinin resistant background. The diffusion of resistant genes in the Cambodian admixed population after selection resulted from mating of gametocytes of sensitive and resistant parasite populations.
In recent years, a number of new protein structures that possess tandem repeats have emerged. Many of these proteins are comprised of tandem arrays of β-hairpins. Today, the amount and variety of the data on these β-hairpin repeat (BHR) structures have reached a level that requires detailed analysis and further classification. In this paper, we classified the BHR proteins, compared structures, sequences of repeat motifs, functions and distribution across the major taxonomic kingdoms of life and within organisms. As a result, we identified six different BHR folds in tandem repeat proteins of Class III (elongated structures) and one BHR fold (up-and-down β-barrel) in Class IV ("closed" structures). Our survey reveals the high incidence of the BHR proteins among bacteria and viruses and their possible relationship to the structures of amyloid fibrils. It indicates that BHR folds will be an attractive target for future structural studies, especially in the context of age-related amyloidosis and emerging infectious diseases. This work allowed us to update the RepeatsDB database, which contains annotated tandem repeat protein structures and to construct sequence profiles based on BHR structural alignments.
Our aim in CASP12 was to improve our Template‐Based Modeling (TBM) methods through better model selection, accuracy self‐estimate (ASE) scores and refinement. To meet this aim, we developed two new automated methods, which we used to score, rank, and improve upon the provided server models. Firstly, the ModFOLD6_rank method, for improved global Quality Assessment (QA), model ranking and the detection of local errors. Secondly, the ReFOLD method for fixing errors through iterative QA guided refinement. For our automated predictions we developed the IntFOLD4‐TS protocol, which integrates the ModFOLD6_rank method for scoring the multiple‐template models that were generated using a number of alternative sequence‐structure alignments. Overall, our selection of top models and ASE scores using ModFOLD6_rank was an improvement on our previous approaches. In addition, it was worthwhile attempting to repair the detected errors in the top selected models using ReFOLD, which gave us an overall gain in performance. According to the assessors' formula, the IntFOLD4 server ranked 3rd/5th (average Z‐score > 0.0/–2.0) on the server only targets, and our manual predictions (McGuffin group) ranked 1st/2nd (average Z‐score > −2.0/0.0) compared to all other groups.
[This corrects the article DOI: 10.1371/journal.pone.0185372.].
Fbw7 is a tumor suppressor often deleted or mutated in human cancers. It serves as the substrate-recruiting subunit of a SCF ubiquitin ligase that targets numerous critical proteins for degradation, including oncoproteins and master transcription factors. Cyclin E was the first identified substrate of the SCFFbw7 ubiquitin ligase. In human cancers bearing FBXW7-gene mutations, deregulation of cyclin E turnover leads to its aberrant expression in mitosis. We investigated Fbw7 regulation in Xenopus eggs, which, although arrested in a mitotic-like phase, naturally express high levels of cyclin E. Here, we report that Fbw7α, the only Fbw7 isoform detected in eggs, is phosphorylated by PKC (protein kinase C) at a key residue (S18) in a manner coincident with Fbw7α inactivation. We show that this PKC-dependent phosphorylation and inactivation of Fbw7α also occurs in mitosis during human somatic cell cycles, and importantly is critical for Fbw7α stabilization itself upon nuclear envelope breakdown. Finally, we provide evidence that S18 phosphorylation, which lies within the intrinsically disordered N-terminal region specific to the α-isoform reduces the capacity of Fbw7α to dimerize and to bind cyclin E. Together, these findings implicate PKC in an evolutionarily-conserved pathway that aims to protect Fbw7α from degradation by keeping it transiently in a resting, inactive state.
Human babesiosis is caused by the apicomplexan parasite Babesia microti, which is of major public health concern in the United States and elsewhere, resulting in malaise and fatigue, followed by a fever and hemolytic anemia. In this paper we focus on the characterization of a novel B. microti thrombospondin domain (TSP1)-containing protein (BmP53) from the new annotation of the B. microti genome (locus 'BmR1_04g09041'). This novel protein (BmP53) had a single TSP1 and a transmembrane domain, with a short cytoplasmic tail containing a sub-terminal glutamine residue, but no signal peptide and Von Willebrand factor type A domains (VWA), which are found in classical thrombospondin-related adhesive proteins (TRAP). Co-localization assays of BmP53 and Babesia microti secreted antigen 1 (BmSA1) suggested that BmP53 might be a non-secretory membranous protein. Molecular mimicry between the TSP1 domain from BmP53 and host platelets molecules was indicated through different measures of sequence homology, phylogenetic analysis, 3D structure and shared epitopes. Indeed, hamster isolated platelets cross-reacted with mouse anti-BmP53-TSP1. Molecular mimicry are used to help parasites to escape immune defenses, resulting in immune evasion or autoimmunity. Furthermore, specific host reactivity was also detected against the TSP1-free part of BmP53 in infected hamster sera. In conclusion, the TSP1 domain mimicry might help in studying the mechanisms of parasite-induced thrombocytopenia, with the TSP1-free truncate of the protein representing a potential safe candidate for future vaccine studies.
There has been an increased interest in computational methods for amyloid and (or) aggregate prediction, due to the prevalence of these aggregates in numerous diseases and their recently discovered functional importance. To evaluate these methods, several datasets have been compiled. Typically, aggregation-prone regions of proteins, which form aggregates or amyloids in vivo, are more than 15 residues long and intrinsically disordered. However, the number of such experimentally established amyloid forming and non-forming sequences are limited, not exceeding one hundred entries in existing databases. In this work, we parsed all available NMR-resolved protein structures from the PDB and assembled a new, sevenfold larger, dataset of unfolded sequences, soluble at high concentrations. We proposed to use these sequences as a negative set for evaluating methods for predicting aggregation in vivo. We also present the results of benchmarking cutting edge tools for the prediction of aggregation versus solubility propensity.
Glycosylphosphatidylinositol (GPI) and GPI-anchored proteins, represent potential new malaria drug targets. GPI proteins are outer plasma membranes anchored proteins, found in the majority of living organisms, including mammals, yeast, protozoan in addition to archaebacteria. Parasitic GPI proteins have numerous functions including; surface coat proteins; receptors; adhesion molecules; enzymes; in addition to possible roles in host-parasite interactions and immune escape. Sensitivity and specificity of machine-learning algorithms predicting GPI-proteins is weakened by the diversity of the N- and C-terminus specific signals. We present a feature-based approach, predicting GPI-proteins from full-length amino-acid sequence. We show that predicting the hydrophobic properties of the N- and C-terminal, in addition to the central core of the protein was more accurate for the prediction of GPI-proteins than current approaches GPI proteins constitute a group which is highly variable from one species to another, but main functions are related to cell-cell interactions. Few enzymatic function are associated with GPI proteins. Therefore, prediction of the GPI proteome can only rely on the detection of the C-terminus region interacting with the enzyme catalyzing transfer of the protein on the GPI anchor.
Damiano Piovesan1,†, Francesco Tabaro1,2,†, Ivan Mičetić1, Marco Necci1, Federica Quaglia1, Christopher J. Oldfield3, Maria Cristina Aspromonte4, Norman E. Davey5,6, Radoslav Davidović7, Zsuzsanna Dosztányi8,9, Arne Elofsson10, Alessandra Gasparini4, András Hatos1,9, Andrey V. Kajava11,12,13, Lajos Kalmar9,14, Emanuela Leonardi4, Tamas Lazar15,16, Sandra Macedo-Ribeiro17, Mauricio Macossay-Castillo15,16, Attila Meszaros9, Giovanni Minervini1, Nikoletta Murvai9, Jordi Pujols18, Daniel B. Roche11,12, Edoardo Salladini19, Eva Schad9, Antoine Schramm19, Beata Szabo9, Agnes Tantos9, Fiorella Tonello1,20, Konstantinos D. Tsirigos10, Nevena Veljković7, Salvador Ventura18, Wim Vranken15,16,21, Per Warholm10, Vladimir N. Uversky22,23, A. Keith Dunker3, Sonia Longhi19,*, Peter Tompa9,15,16,* and Silvio C.E. Tosatto1,20,*
Protein tertiary structure prediction algorithms aim to predict, from amino acid sequence, the tertiary structure of a protein. In silico protein structure prediction methods have become extremely important, as in vitro-based structural elucidation is unable to keep pace with the current growth of sequence databases due to high-throughput next-generation sequencing, which has exacerbated the gaps in our knowledge between sequences and structures.Here we briefly discuss protein tertiary structure prediction, the biennial competition for the Critical Assessment of Techniques for Protein Structure Prediction (CASP) and its role in shaping the field. We also discuss, in detail, our cutting-edge web-server method IntFOLD2-TS for tertiary structure prediction. Furthermore, we provide a step-by-step guide on using the IntFOLD2-TS web server, along with some real world examples, where the IntFOLD server can and has been used to improve protein tertiary structure prediction and aid in functional elucidation.
Despite the ever-increasing role of pesticides in modern agriculture, their deleterious effects are still underexplored. Here we examine the effect of A6, a pesticide derived from the naturally-occurring α-terthienyl, and structurally related to the endocrine disrupting pesticides anilinopyrimidines, on living zebrafish larvae. We show that both A6 and an anilinopyrimidine, cyprodinyl, decrease larval survival and affect central neurons at micromolar concentrations. Focusing on a superficial and easily observable sensory system, the lateral line system, we found that defects in axonal and sensory cell regeneration can be observed at much lower doses, in the nanomolar range. We also show that A6 accumulates preferentially in lateral line neurons and hair cells. We examined whether A6 affects the expression of putative target genes, and found that genes involved in apoptosis/cell proliferation are down-regulated, as well as genes reflecting estrogen receptor activation, consistent with previous reports that anilinopyrimidines act as endocrine disruptors. On the other hand, canonical targets of endocrine signaling are not affected, suggesting that the neurotoxic effect of A6 may be due to the binding of this compound to a recently identified, neuron-specific estrogen receptor.
Protein-ligand binding site prediction methods aim to predict, from amino acid sequence, protein-ligand interactions, putative ligands, and ligand binding site residues using either sequence information, structural information, or a combination of both. In silico characterization of protein-ligand interactions has become extremely important to help determine a protein's functionality, as in vivo-based functional elucidation is unable to keep pace with the current growth of sequence databases. Additionally, in vitro biochemical functional elucidation is time-consuming, costly, and may not be feasible for large-scale analysis, such as drug discovery. Thus, in silico prediction of protein-ligand interactions must be utilized to aid in functional elucidation. Here, we briefly discuss protein function prediction, prediction of protein-ligand interactions, the Critical Assessment of Techniques for Protein Structure Prediction (CASP) and the Continuous Automated EvaluatiOn (CAMEO) competitions, along with their role in shaping the field. We also discuss, in detail, our cutting-edge web-server method, FunFOLD for the structurally informed prediction of protein-ligand interactions. Furthermore, we provide a step-by-step guide on using the FunFOLD web server and FunFOLD3 downloadable application, along with some real world examples, where the FunFOLD methods have been used to aid functional elucidation.
The Database of Protein Disorder (DisProt, URL: www.disprot.org) has been significantly updated and upgraded since its last major renewal in 2007. The current release holds information on more than 800 entries of IDPs/IDRs, i.e. intrinsically disordered proteins or regions that exist and function without a well-defined three-dimensional structure. We have re-curated previous entries to purge DisProt from conflicting cases, and also upgraded the functional classification scheme to reflect continuous advance in the field in the past 10 years or so. We define IDPs as proteins that are disordered along their entire sequence, i.e. entirely lack structural elements, and IDRs as regions that are at least five consecutive residues without well-defined structure. We base our assessment of disorder strictly on experimental evidence, such as X-ray crystallography and nuclear magnetic resonance ( primary techniques) and a broad range of other experimental approaches (secondary techniques). Confident and ambiguous annotations are highlighted separately. DisProt 7.0 presents classified knowledge regarding the experimental characterization and functional annotations of IDPs/IDRs, and is intended to provide an invaluable resource for the research community for a better understanding structural disorder and for developing better computational tools for studying disordered proteins.
IntFOLD is an independent web server that integrates our leading methods for structure and function prediction. The server provides a simple unified interface that aims to make complex protein modelling data more accessible to life scientists. The server web interface is designed to be intuitive and integrates a complex set of quantitative data, so that 3D modelling results can be viewed on a single page and interpreted by non-expert modellers at a glance. The only required input to the server is an amino acid sequence for the target protein. Here we describe major performance and user interface updates to the server, which comprises an integrated pipeline of methods for: tertiary structure prediction, global and local 3D model quality assessment, disorder prediction, structural domain prediction, function prediction and modelling of protein-ligand interactions. The server has been independently validated during numerous CASP (Critical Assessment of Techniques for Protein Structure Prediction) experiments, as well as being continuously evaluated by the CAMEO (Continuous Automated Model Evaluation) project. The IntFOLD server is available at: http://www.reading.ac.uk/bioinf/IntFOLD/.
As the largest fraction of any proteome does not carry out enzymatic functions, and in order to leverage 3D structural data for the annotation of increasingly higher volumes of sequence data, we wanted to assess the strength of the link between coarse grained structural data (i.e., homologous superfamily level) and the enzymatic versus non-enzymatic nature of protein sequences. To probe this relationship, we took advantage of 41 phylogenetically diverse (encompassing 11 distinct phyla) genomes recently sequenced within the GEBA initiative, for which we integrated structural information, as defined by CATH, with enzyme level information, as defined by Enzyme Commission (EC) numbers. This analysis revealed that only a very small fraction (about 1%) of domain sequences occurring in the analyzed genomes was found to be associated with homologous superfamilies strongly indicative of enzymatic function. Resorting to less stringent criteria to define enzyme versus non-enzyme biased structural classes or excluding highly prevalent folds from the analysis had only modest effect on this proportion. Thus, the low genomic coverage by structurally anchored protein domains strongly associated to catalytic activities indicates that, on its own, the power of coarse grained structural information to infer the general property of being an enzyme is rather limited.
Elucidating the biological and biochemical roles of proteins, and subsequently determining their interacting partners, can be difficult and time consuming using in vitro and/or in vivo methods, and consequently the majority of newly sequenced proteins will have unknown structures and functions. However, in silico methods for predicting protein–ligand binding sites and protein biochemical functions offer an alternative practical solution. The characterisation of protein–ligand binding sites is essential for investigating new functional roles, which can impact the major biological research spheres of health, food, and energy security. In this review we discuss the role in silico methods play in 3D modelling of protein–ligand binding sites, along with their role in predicting biochemical functionality. In addition, we describe in detail some of the key alternative in silico prediction approaches that are available, as well as discussing the Critical Assessment of Techniques for Protein Structure Prediction (CASP) and the Continuous Automated Model EvaluatiOn (CAMEO) projects, and their impact on developments in the field. Furthermore, we discuss the importance of protein function prediction methods for tackling 21st century problems.