Rare diseases, though cumulatively common, can be complex to diagnose individually. It is becoming standard practice to sequence and analyse exomes or genomes to diagnose affected individuals, and gain insight into possible pathophysiological mechanisms. An improved understanding of the impact of genetic variation is required to accelerate this process. Mapping variants to genes and predicting their effects at the protein level is a key first step in interpretation. Knowledge of previously discovered links between phenotypes and both variants and the genes they lie within is essential. Simple extraction of information showing the pathways genes are involved in can provide insight into potential disease mechanisms. Mature tools are available to enable this, including: the Ensembl Variant Effect Predictor, which annotates variants with predicted gene consequence, pathogenicity scores and population frequencies; neXtProt, a reference database on human proteins which provides detailed information about the functional impact of mutations for a set of clinically relevant proteins; Orphanet which curates gene-rare disease associations; the Leiden Open Variation Database (LOVD) and ClinVar which capture variant phenotype associations; and WikiPathways, the community-curated database for biological pathways. As part of the European Joint Project on Rare Disease, we seek to improve interoperability and to facilitate federated queries across existing resources. We will also enhance the functionality of our tools through discussions within the community to elicit requirements. This work will contribute to the construction of the ELIXIR service bundles for Rare diseases.
The neXtProt platform ( www nextprot org ), developed at the SIB Swiss Institute of Bioinformatics the Swiss ELIXIR node, is a one stop shop for human proteins. neXtProt proposes solutions to explore, select and reuse available genomic, transcriptomic, mass spectrometry and antibody based proteomic data. The neXtProt team manually curates data from the literature (post translational modifications, variant phenotypes, protein protein interactions, etc and combines it with high quality omics data generated by systems biology projects using a single inter operable format. neXtProt data are FAIR ( Accessible, Interoperable, and Reusable), with full traceability ensured by extensive use of metadata. In the last four years, neXtProt has been promoting the use of SPARQL, a semantic query language for databases, to explore human data. Semantic technologies can help to generate innovative hypotheses where classical data mining tools have failed (protein function prediction, drug repositioning, etc.). In order to promote the use of semantic technologies as data mining tools for life sciences, neXtProt provides over 140 pre built queries and documentation of its data model to guide the user in his/her first steps. The use of SPARQL allows users to run federated queries across multiple resources relevant to human biology.
The neXtProt knowledgebase (https://www.nextprot.org) is an integrative resource providing both data on human protein and the tools to explore these. In order to provide comprehensive and up-to-date data, we evaluate and add new data sets. We describe the incorporation of three new data sets that provide expression, function, protein-protein binary interaction, post-translational modifications (PTM) and variant information. New SPARQL query examples illustrating uses of the new data were added. neXtProt has continued to develop tools for proteomics. We have improved the peptide uniqueness checker and have implemented a new protein digestion tool. Together, these tools make it possible to determine which proteases can be used to identify trypsin-resistant proteins by mass spectrometry. In terms of usability, we have finished revamping our web interface and completely rewritten our API. Our SPARQL endpoint now supports federated queries. All the neXtProt data are available via our user interface, API, SPARQL endpoint and FTP site, including the new PEFF 1.0 format files. Finally, the data on our FTP site is now CC BY 4.0 to promote its reuse.
The neXtProt platform ( www.nextprot.org ) developed at SIB Swiss Institute of Bioinformatics is a one-stop-shop for human proteins proposing solutions to select, explore and reuse available genomic, transcriptomics, mass-spectrometry- and antibody-based proteomics data. The neXtProt team manually curates data from the literature (posttranslational modifications, variant phenotypes, protein-protein interactions, etc.) and combines it with high quality omics data generated by systems biology projects using a single inter-operable format. neXtProt data are FAIR (Findable, Accessible, Interoperable, and Reusable), with full traceability ensured by extensive use of metadata. In the last four years, neXtProt has been promoting the use of SPARQL, a semantic query language for databases, to check, explore, and visualize its data. SPARQL queries are used to check the quality and consistency of the data loaded at each release. To date, over 450 queries have been written such that non-zero results trigger investigation. In an effort to automate these tests, all of the queries or a particular sub-set can be launched and the results written to a file. Semantic technologies can help generating innovative hypotheses where classical data mining tools have failed (protein function prediction, drug repositioning...). In order to promote the use of semantic technologies as data mining tools for life sciences, neXtProt provides over 140 pre-built queries and documentation of its data model to guide the user in his or her first steps. The use of SPARQL allows users to run federated queries across resources relevant for human biology, or build customized views the data. All our SPARQL queries are open source and available on GitHub ( github.com/calipho-sib/nextprot-queries ).
20,230 protein-coding genes have been predicted from the analysis of the human genome (neXtProt release 2018-01-17), and about 10% of them are still lacking functional annotation, either predicted by bioinformatics tools or captured from experimental reports. A systematic exploration of the available literature on uncharacterized human genes/proteins led to proposal of functional annotations for 113 proteins and to consolidation of a list of 1,862 uncharacterized human proteins. The advanced search functionality of neXtProt was used extensively in order to examine the landscape of the uncharacterized human proteome in terms of subcellular locations, protein-protein interactions, tissue expression, association with diseases, and 3D structure. Finally, a deep data mining in various publicly available resources allowed building functional hypotheses for 26 uncharacterized human proteins validated at protein level (uPE1). These hypotheses cover the fields of cilia biology, male reproduction, metabolism, nervous system, immunity, inflammation, RNA metabolism, and chromatin biology. They will require experimental validation before they can be considered for annotation. Despite technological progresses, the pace of human protein characterization studies is still slow. It could be accelerated by a better integration of existing knowledge resources and by initiating large collaborative projects involving specialists of different biology fields. We hope that our analysis will contribute to set up the ground for such collaborative approaches and will be exploited by the HUPO Human Proteome Project teams committed to characterize uPE1 proteins.
Abstract Summary The neXtProt peptide uniqueness checker allows scientists to define which peptides can be used to validate the existence of human proteins, i.e. map uniquely versus multiply to human protein sequences taking into account isobaric substitutions, alternative splicing and single amino acid variants. Availability and implementation The pepx program is available at https://github.com/calipho-sib/pepx and can be launched from the command line or through a cgi web interface. Indexing requires a sequence file in FASTA format. The peptide uniqueness checker tool is freely available on the web at https://www.nextprot.org/tools/peptide-uniqueness-checker and from the neXtProt API at https://api.nextprot.org/.
The neXtProt human protein knowledgebase (https://www.nextprot.org) continues to add new content and tools, with a focus on proteomics and genetic variation data. neXtProt now has proteomics data for over 85% of the human proteins, as well as new tools tailored to the proteomics community.Moreover, the neXtProt release 2016-08-25 includes over 8000 phenotypic observations for over 4000 variations in a number of genes involved in hereditary cancers and channelopathies. These changes are presented in the current neXtProt update. All of the neXtProt data are available via our user interface and FTP site. We also provide an API access and a SPARQL endpoint for more technical applications.
The SIB Swiss Institute of Bioinformatics (www.isb-sib.ch) provides world-class bioinformatics databases, software tools, services and training to the international life science community in academia and industry. These solutions allow life scientists to turn the exponentially growing amount of data into knowledge. Here, we provide an overview of SIB's resources and competence areas, with a strong focus on curated databases and SIB's most popular and widely used resources. In particular, SIB's Bioinformatics resource portal ExPASy features over 150 resources, including UniProtKB/Swiss-Prot, ENZYME, PROSITE, neXtProt, STRING, UniCarbKB, SugarBindDB, SwissRegulon, EPD, arrayMap, Bgee, SWISS-MODEL Repository, OMA, OrthoDB and other databases, which are briefly described in this article.
Within the C-HPP, the Swiss and French teams are responsible for the annotation of proteins from chromosomes 2 and 14, respectively. neXtProt currently reports 1231 entries on chromosome 2 and 624 entries on chromosome 14; of these, 134 and 93 entries are still not experimentally validated and are thus considered as "missing proteins" (PE2-4), respectively. Among these entries, some may never be validated by conventional MS/MS approaches because of incompatible biochemical features. Others have already been validated but are still awaiting annotation. On the basis of information retrieved from the literature and from three of the main C-HPP resources (Human Protein Atlas, PeptideAtlas, and neXtProt), a subset of 40 theoretically detectable missing proteins (25 on chromosome 2 and 15 on chromosome 14) was defined for upcoming targeted studies in sperm samples. This list is proposed as a roadmap for the French and Swiss teams in the near future.
The Chromosome-Centric Human Proteome Project (C-HPP) aims to identify "missing" proteins in the neXtProt knowledgebase. We present an in-depth proteomics analysis of the human sperm proteome to identify testis-enriched missing proteins. Using protein extraction procedures and LC-MS/MS analysis, we detected 235 proteins (PE2-PE4) for which no previous evidence of protein expression was annotated. Through LC-MS/MS and LC-PRM analysis, data mining, and immunohistochemistry, we confirmed the expression of 206 missing proteins (PE2-PE4) in line with current HPP guidelines (version 2.0). Parallel reaction monitoring acquisition and sythetic heavy labeled peptides targeted 36 ≪one-hit wonder≫ candidates selected based on prior peptide spectrum match assessment. 24 were validated with additional predicted and specifically targeted peptides. Evidence was found for 16 more missing proteins using immunohistochemistry on human testis sections. The expression pattern for some of these proteins was specific to the testis, and they could possibly be valuable markers with fertility assessment applications. Strong evidence was also found of four "uncertain" proteins (PE5); their status should be re-examined. We show how using a range of sample preparation techniques combined with MS-based analysis, expert knowledge, and complementary antibody-based techniques can produce data of interest to the community. All MS/MS data are available via ProteomeXchange under identifier PXD003947. In addition to contributing to the C-HPP, we hope these data will stimulate continued exploration of the sperm proteome.
In the framework of the C-HPP, our Franco-Swiss consortium has adopted chromosomes 2 and 14, coding for a total of 382 missing proteins (proteins for which evidence is lacking at protein level). Over the last 4 years, the French proteomics infrastructure has collected high-quality data sets from 40 human samples, including a series of rarely studied cell lines, tissue types, and sample preparations. Here we described a step-by-step strategy based on the use of bioinformatics screening and subsequent mass spectrometry (MS)-based validation to identify what were up to now missing proteins in these data sets. Screening database search results (85,326 dat files) identified 58 of the missing proteins (36 on chromosome 2 and 22 on chromosome 14) by 83 unique peptides following the latest release of neXtProt (2014-09-19). PSMs corresponding to these peptides were thoroughly examined by applying two different MS-based criteria: peptide-level false discovery rate calculation and expert PSM quality assessment. Synthetic peptides were then produced and used to generate reference MS/MS spectra. A spectral similarity score was then calculated for each pair of reference-endogenous spectra and used as a third criterion for missing protein validation. Finally, LC-SRM assays were developed to target proteotypic peptides from four of the missing proteins detected in tissue/cell samples, which were still available and for which sample preparation could be reproduced. These LC-SRM assays unambiguously detected the endogenous unique peptide for three of the proteins. For two of these, identification was confirmed by additional proteotypic peptides. We concluded that of the initial set of 58 proteins detected by the bioinformatics screen, the consecutive MS-based validation criteria led to propose the identification of 13 of these proteins (8 on chromosome 2 and 5 on chromosome 14) that passed at least two of the three MS-based criteria. Thus, a rigorous step-by-step approach combining bioinformatics screening and MS-based validation assays is particularly suitable to obtain protein-level evidence for proteins previously considered as missing. All MS/MS data have been deposited in ProteomeXchange under identifier PXD002131.
The Chromosome-Centric Human Proteome Project (C-HPP) aims at cataloguing the proteins as gene products encoded by the human genome in a chromosome-centric manner. The existence of products of about 82% of the genes has been confirmed at the protein level. However, the number of so-called "missing proteins" remains significant. It was recently suggested that the expression of proteins that have been systematically missed might be restricted to particular organs or cell types, for example, the testis. Testicular function, and spermatogenesis in particular, is conditioned by the successive activation or repression of thousands of genes and proteins including numerous germ cell- and testis-specific products. Both the testis and postmeiotic germ cells are thus promising sites at which to search for missing proteins, and ejaculated spermatozoa are a potential source of proteins whose expression is restricted to the germ cell lineage. A trans-chromosome-based data analysis was performed to catalog missing proteins in total protein extracts from isolated human spermatozoa. We have identified and manually validated peptide matches to 89 missing proteins in human spermatozoa. In addition, we carefully validated three proteins that were scored as uncertain in the latest neXtProt release (09.19.2014). A focus was then given to the 12 missing proteins encoded on chromosomes 2 and 14, some of which may putatively play roles in ciliation and flagellum mechanistics. The expression pattern of C2orf57 and TEX37 was confirmed in the adult testis by immunohistochemistry. On the basis of transcript expression during human spermatogenesis, we further consider the potential for discovering additional missing proteins in the testicular postmeiotic germ cell lineage and in ejaculated spermatozoa. This project was conducted as part of the C-HPP initiatives on chromosomes 14 (France) and 2 (Switzerland). The mass spectrometry proteomics data have been deposited with the ProteomeXchange Consortium under the data set identifier PXD002367.
About 5000 (25%) of the ~20400 human protein-coding genes currently lack any experimental evidence at the protein level. For many others, there is only little information relative to their abundance, distribution, subcellular localization, interactions, or cellular functions. The aim of the HUPO Human Proteome Project (HPP, www.thehpp.org ) is to collect this information for every human protein. HPP is based on three major pillars: mass spectrometry (MS), antibody/affinity capture reagents (Ab), and bioinformatics-driven knowledge base (KB). To meet this objective, the Chromosome-Centric Human Proteome Project (C-HPP) proposes to build this catalog chromosome-by-chromosome ( www.c-hpp.org ) by focusing primarily on proteins that currently lack MS evidence or Ab detection. These are termed "missing proteins" by the HPP consortium. The lack of observation of a protein can be due to various factors including incorrect and incomplete gene annotation, low or restricted expression, or instability. neXtProt ( www.nextprot.org ) is a new web-based knowledge platform specific for human proteins that aims to complement UniProtKB/Swiss-Prot ( www.uniprot.org ) with detailed information obtained from carefully selected high-throughput experiments on genomic variation, post-translational modifications, as well as protein expression in tissues and cells. This article describes how neXtProt contributes to prioritize C-HPP efforts and integrates C-HPP results with other research efforts to create a complete human proteome catalog.
The primary mission of Universal Protein Resource (UniProt) is to support biological research by maintaining a stable, comprehensive, fully classified, richly and accurately annotated protein sequence knowledgebase, with extensive cross-references and querying interfaces freely accessible to the scientific community. UniProt is produced by the UniProt Consortium which consists of groups from the European Bioinformatics Institute (EBI), the Swiss Institute of Bioinformatics (SIB) and the Protein Information Resource (PIR). UniProt is comprised of four major components, each optimized for different uses: the UniProt Archive, the UniProt Knowledgebase, the UniProt Reference Clusters and the UniProt Metagenomic and Environmental Sequence Database. UniProt is updated and distributed every 4 weeks and can be accessed online for searches or download at http://www.uniprot.org.
neXtProt http://www.nextprot.org/ is a new human protein-centric knowledge platform. Developed at the Swiss Institute of Bioinformatics (SIB), it aims to help researchers answer questions relevant to human proteins. To achieve this goal, neXtProt is built on a corpus containing both curated knowledge originating from the UniProtKB/Swiss-Prot knowledgebase and carefully selected and filtered high-throughput data pertinent to human proteins. This article presents an overview of the database and the data integration process. We also lay out the key future directions of neXtProt that we consider the necessary steps to make neXtProt the one-stop-shop for all research projects focusing on human proteins.
The primary mission of UniProt is to support biological research by maintaining a stable, comprehensive, fully classified, richly and accurately annotated protein sequence knowledgebase, with extensive cross-references and querying interfaces freely accessible to the scientific community. UniProt is produced by the UniProt Consortium which consists of groups from the European Bioinformatics Institute (EBI), the Swiss Institute of Bioinformatics (SIB) and the Protein Information Resource (PIR). UniProt is comprised of four major components, each optimized for different uses: the UniProt Archive, the UniProt Knowledgebase, the UniProt Reference Clusters and the UniProt Metagenomic and Environmental Sequence Database. UniProt is updated and distributed every 3 weeks and can be accessed online for searches or download at http://www.uniprot.org.
neXtProt (http://www.nextprot.org) is a human protein-centric knowledgebase developed at the SIB Swiss Institute of Bioinformatics. Focused solely on human proteins, neXtProt aims to provide a state of the art resource for the representation of human biology by capturing a wide range of data, precise annotations, fully traceable data provenance and a web interface which enables researchers to find and view information in a comprehensive manner. Since the introductory neXtProt publication, significant advances have been made on three main aspects: the representation of proteomics data, an extended representation of human variants and the development of an advanced search capability built around semantic technologies. These changes are presented in the current neXtProt update.
The primary mission of UniProt is to support biological research by maintaining a stable, comprehensive, fully classified, richly and accurately annotated protein sequence knowledgebase, with extensive cross-references to external resources, that is freely available to the scientific community. To enable users of the knowledgebase to accurately assess the reliability of the information contained in this resource, the evidence for and provenance of the information must be recorded. This paper discusses the user requirements for this kind of metadata and the manner in which UniProtKB records it.
The mission of UniProt is to provide the scientific community with a comprehensive, high-quality and freely accessible resource of protein sequence and functional information that is essential for modern biological research. UniProt is produced by the UniProt Consortium which consists of groups from the European Bioinformatics Institute, the Protein Information Resource and the Swiss Institute of Bioinformatics. The core activities include manual curation of protein sequences assisted by computational analysis, sequence archiving, a user-friendly UniProt website and the provision of additional value-added information through cross-references to other databases. UniProt is comprised of four major components, each optimized for different uses: the UniProt Archive, the UniProt Knowledge-base, the UniProt Reference Clusters and the UniProt Metagenomic and Environmental Sequence Database. One of the key achievements of the UniProt consortium in 2008 is the completion of the first draft of the complete human proteome in UniProtKB/Swiss-Prot. This manually annotated representation of all currently known human protein-coding genes was made available in UniProt release 14.0 with 20 325 entries. UniProt is updated and distributed every three weeks and can be accessed online for searches or downloaded at www.uniprot.org.
UniProtKB/Swiss-Prot (http://beta.uniprot.org/uniprot; last accessed: 19 October 2007) is a manually curated knowledgebase providing information on protein sequences and functional annotation. It is part of the Universal Protein Resource (UniProt). The knowledgebase currently records a total of 32,282 single amino acid polymorphisms (SAPs) touching 6,086 human proteins (Release 53.2, 26 June 2007). Nearly all SAPs are derived from literature reports using strict inclusion criteria. For each SAP, the knowledgebase provides, apart from the position of the mutation and the resulting change in amino acid, information on the effects of SAPs on protein structure and function, as well as their potential involvement in diseases. Presently, there are 16,043 disease-related SAPs, 14,266 polymorphisms, and 1,973 unclassified variants recorded in UniProtKB/Swiss-Prot. Relevant information on SAPs can be found in various sections of a UniProtKB/Swiss-Prot entry. In addition to these, cross-references to human disease databases as well as other gene-specific databases, are being added regularly. In 2003, the Swiss-Prot variant pages were created to provide a concise view of the information related to the SAPs recorded in the knowledgebase. When compared to the information on missense variants listed in other mutation databases, UniProtKB/Swiss-Prot further records information on direct protein sequencing and characterization including posttranslational modifications (PTMs). The direct links to the Online Mendelian Inheritance in Man (OMIM) database entries further enhance the integration of phenotype information with data at protein level. In this regard, SAP information in UniProtKB/Swiss-Prot complements nicely those existing in genomic and phenotypic databases, and is valuable for the understanding of SAPs and diseases.
Brigitte Boeckmann合作论文数Swiss Institute of Bioinformatics6