Background The discovery of autoantibodies against citrullinated proteins (ACPA) and the consequent development of serological tests using cyclic citrullinated peptide (CCP) have improved diagnosis of rheumatoid arthritis (RA). ACPAs recognize specifically various citrullinated peptides that derive from different antigens. It has also been demonstrated that autoantibodies to different citrullinated peptides emerge at different stages of RA. However, by now the fine specificity against only a limited number of different peptides has been tested. Objectives To increase knowledge about ACPAs and their fine specificities in RA patients we aimed to identify novel citrullinated antigens and to characterize the ACPA profile in early and late RA. Methods 569 serum samples of CCP+ and CCP- RA patients with late RA, early RA or pre-RA and additional samples of volunteers and patients with early Arthritis (EA) but without RA at the time point tested have been analyzed in a multiplexed, bead-based, antigen array on the Luminex FlexMAP3D. Sera have been tested against a selection of 417 human proteins, including previously identified putative RA specific autoantigens and also known RA marker such as vimentin and fibrinogen. After coupling all proteins to individual color coded Luminex beads they were enzymatically citrullinated by peptidyl arginine deiminase and then tested against sera in the multiplex Luminex array. Complementarily all proteins were also measured in their native, unmodified form. Results We found novel citrullinated antigens targeted by ACPAs. The citrullinated antigens are highly reactive in mature RA whereas the unmodified proteins show only low reactivity. Generally, we observe an increase of the number of ACPA fine specificities with disease duration. In CCP negative tested RA patients, more than 10% of the patients recognize specifically several citrullinated antigens suggesting their diagnostic potential. Conclusions Novel citrullinated antigens targeted by ACPAs were identified. The identification of ACPA positive but CCP negative tested RA patients suggests that current CCP tests do not address the complete profile of ACPA fine specificities. Further studies are currently conducted to validate panels of citrullinated antigens in early RA patients. Disclosure of Interest A. Lueking Employee of: Protagen AG, P. Schriek Employee of: Protagen AG, H. Goehler Employee of: Protagen AG, M. Gamer Employee of: Protagen AG, K. Marquart Employee of: Protagen AG, A. Telaar: None declared, D. Chamrad Employee of: Protagen AG, P. Schulz-Knappe Employee of: Protagen AG, J. Richter: None declared, S. Vordenbaeumen: None declared, G. Burmester: None declared, M. Schneider: None declared
Background Systemic lupus erythematosus (SLE) is a multifactorial autoimmune disease with great heterogeneity between affected individuals. This leads to difficulties in predicting disease manifestations and developing effective new therapies for SLE. Several immune defects have been described, which lead to the overproduction of autoantibodies against cellular and nuclear antigens and immune-mediated damage of several organs and tissues. The in depth-characterization of the autoantibody repertoire in SLE is an emerging tool to identify biomarkers supporting a personalized disease management approach. We have recently conducted high-content autoantibody profiling studies of SLE, systemic autoimmune diseases (AID), and healthy controls and found in addition to diagnostic autoantibodies novel SLE-associated autoantibodies (1). Objectives The multiplex bead-based Luminex® xMAP® technology was applied to define a set of autoantigens that were consistently measured in the discovery and validation cohort. Confirmed antigens were employed to identify clustered autoantibodies and their clinical association in SLE. Methods The discovery phase constituted the broad characterization of serum samples from 130 SLE patients, 794 AID patients (systemic sclerosis, rheumatoid arthritis/RA, early RA, ankylosing spondylitis) and 343 healthy controls. Autoantibody reactivity was tested against 6,912 recombinant human proteins using the bead-based Luminex technology. SLE-specific autoantibody reactivity (p-value <0.05, effect size: Cohen9s d >0.3) was observed for 166 antigens, consisting of well-known plus novel AID marker candidates. Serum samples from the SLE discovery cohort and a new SLE cohort (n=101) were then used for technical verification and validation of the markers. Results Based on autoantibody discovery screens, a list of 166 candidate SLE marker antigens was generated. Following technical replication and validation in independent samples, consistent autoantibody reactivity against 46 antigens was found (p-value <0.05 and Cohen9s d effect size <0.3. The final list includes among others known SLE antigens (Ro52/SS-A, ribosomal proteins and components of the U1-RNP complex). Novel antigens comprise proteins that localize to the nucleus, cytoplasm, mitochondria and plasma membrane. The top marker were combined into biomarker panels, which demonstrated a sensitivity of 79% and specificity of 81% in the discovery cohort and a sensitivity of 74% and specificity of 82% in the validation cohort. Furthermore, clustered autoantibody reactivity cluster was observed, which can be employed to define subgroups of SLE patients based on their autoantibody profile. Conclusions Using a multiplex platform we have discovered novel SLE-associated antigens. The reproducibility of the autoantibody reactivity was verified and validated in new SLE samples applying a stepwise marker refinement approach. This approach yielded novel markers combinations, which improve the diagnostic sensitivity and specificity. The heterogenic and clustered autoantibody reactivity pattern contributes to defining subgroups of SLE patients for personalized disease management approaches. References Lueking A. et al. Ann Rheum Dis 2013;72:A535. Disclosure of Interest P. Budde Employee of: Protagen AG, S. Vordenbäumen: None declared, D. Chamrad Employee of: Protagen AG, L. Schlieker Employee of: Protagen AG, A. Telaar Employee of: Protagen AG, H.-D. Zucht Employee of: Protagen AG, P. Schulz-Knappe Employee of: Protagen AG, M. Schneider: None declared
Background SLE is a chronic multisystem autoimmune disease with yet unknown etiology. One hallmark of SLE is the B cell activation and the production of various autoantibodies. Differential diagnosis involves ANA screening which may yield also positive results in many connective tissue disorders and other autoimmune diseases, and may occur in normal individuals. More specific subtypes of antinuclear antibodies include anti-Smith and anti-double stranded DNA (dsDNA) antibodies. However, anti-dsDNA antibodies are present in just 70% SLE patients leaving a certain gap in diagnosis. Objectives In this study we characterized the autoimmune signatures of SLE patients and healthy blood donors in order to identify established plus novel autoantibodies with differential abundance in serum samples of SLE patients. Methods We used 5800 human proteins to investigate the autoantibodies present in more than 110 serum samples of SLE patients as well as healthy blood donors applying bead-based arrays (Luminex). Biostatistical analysis was performed using univariate as well as multivariate statistical algorithms. Results In addition to the well-established clinical markers Ro52/SS-A (TRIM21), ribosomal P0, ribosomal P1, ribosomal P2, SS-A/Ro60 (TROVE2), SSB and SmB/B´ new biomarkers were identified showing high discriminatory p-values and promising sensitivity and specificity. Consequently, a small protein antigen panel consisting of four new markers plus six well-established clinical markers (Ro52/SS-A, Ro60/SS-A, ribosomal P0, ribosomal P1, ribosomal P2, and SmB/B´protein) has been defined. This panel results in an improved AUC when compared to the classification performance of the established markers alone. Conclusions With novel biomarkers a large proportion of dsDNA AB-negative SLE patients can be addressed. The benefit of the identified marker panel for SLE diagnosis is currently evaluated in comparative analysis of other patient cohorts of ankylosing spondylitis, progressive systemic sclerosis, and rheumatoid arthritis. Disclosure of Interest None Declared
Within this chapter, various techniques and instructions for characterizing primary structure of proteins are presented, whereas the focus lies on obtaining as much complete sequence information of single proteins as possible. Especially, in the area of protein production, mass spectrometry-based detailed protein characterization plays an increasing important role for quality control. In comparison to typical proteomics applications, wherein it is mostly sufficient to identify proteins by few peptides, several complementary techniques have to be applied to maximize primary structure information and analysis steps have to be specifically adopted. Starting from sample preparation down to mass spectrometry analysis and finally to data analysis, some of the techniques typically applied are outlined here in a summarizing and introductory manner.
MOTIVATIONProteomics has particularly evolved to become of high interest for the field of biomarker discovery and drug development. Especially the combination of liquid chromatography and mass spectrometry (LC/MS) has proven to be a powerful technique for analyzing protein mixtures. Clinically orientated proteomic studies will have to compare hundreds of LC/MS runs at a time. In order to compare different runs, sophisticated preprocessing steps have to be performed. An important step is the retention time (rt) alignment of LC/MS runs. Especially non-linear shifts in the rt between pairs of LC/MS runs make this a crucial and non-trivial problem.RESULTSFor the purpose of demonstrating the particular importance of correcting non-linear rt shifts, we evaluate and compare different alignment algorithms. We present and analyze two versions of a new algorithm that is based on regression techniques, once assuming and estimating only linear shifts and once also allowing for the estimation of non-linear shifts. As an example for another type of alignment method we use an established alignment algorithm based on shifting vectors that we adapted to allow for correcting non-linear shifts also. In a simulation study, we show that rt alignment procedures that can estimate non-linear shifts yield clearly better alignments. This is even true under mild non-linear deviations.AVAILABILITYR code for the regression-based alignment methods and simulated datasets are available at http://www.statistik.tu-dortmund.de/genetik-publikationen-alignment.html.SUPPLEMENTARY INFORMATIONSupplementary data are available at Bioinformatics online.
one of the major challenges for large scale proteomics research is the quality evaluation of results. Protein identification from complex biological samples or experimental setups is often a manual and subjective task which lacks profound statistical evaluation. This is not feasible for high-throughput proteomic experiments which result in large datasets of thousands of peptides and proteins and their corresponding mass spectra. To improve the quality, reliability and comparability of scientific results, an estimation of the rate of erroneously identified proteins is advisable. Moreover, scientific journals increasingly stipulate that articles containing considerable MS data should be subject to stringent statistical evaluation. We present a newly developed easy-to-use software tool enabling quality evaluation by generating composite target-decoy databases usable with all relevant protein search engines. This tool, when used in conjunction with relevant statistical quality criteria, enables to reliably determine peptides and proteins of high quality, even for nonexperienced users (e.g. laboratory staff, researchers without programming knowledge). Different strategies for building decoy databases are implemented and the resulting databases are characterized and compared. The quality of protein identification in high-throughput proteomics is usually measured by the false positive rate (FPR), but it is shown that the false discovery rate (FDR) delivers a more meaningful, robust and comparable value.
Proteomics inherently deals with huge amounts of data. Current mass spectrometers acquire hundreds of thousands of spectra within a single project. Thus, data management and data analysis are a challenge. We have developed a software platform (Proteinscape) that stores all relevant proteomics data efficiently and allows fast access and correlation analysis within proteomics projects. The software is based on a relational database system using Web-based server-client architecture with intra- and Internet access. Proteinscape stores relevant data from all steps of proteomics projects—study design, sample treatment, separation techniques (e.g., gel electrophoresis or liquid chromatography), protein digestion, mass spectrometry, and protein database search results. Gel spot data can be imported directly from several 2DE-gel image analysis software packages as well as spot-picking robots. Spectra (MS and MS/MS) are imported automatically during acquisition from MALDI and ESI mass spectrometers. Many algorithms for automated spectra and search result processing are integrated. PMF spectra are calibrated and filtered for contaminant and polymer peaks (Score-booster). A single non-redundant protein list—containing only proteins that can be distinguished by the MS/MS data—can be generated from MS/MS search results (ProteinExtractor). This algorithm can combine data from different search algorithms or different experiments (MALDI/ESI, or acquisition repetitions) into a single protein list. Navigation within the database is possible either by using the hierarchy of project, sample, protein/peptide separation, spectrum, and identification results, or by using a gel viewer plug-in. Available features include zooming, annotations (protein, spot name, etc.), export of the annotated image, and links to spot, spectrum, and protein data. Proteinscape includes sophisticated query tools that allow data retrieval for typical questions in proteome projects. Here we present the benefit and power of usage of 6 years of continuous use of the software in over 70 proteome projects managed in house.
Available peptide fragmentation interpretation software is focused on sequence database–driven protein identification, rather than on primary structure elucidation. Searching for posttranslational modification (PTM) or sequence errors currently needs time-critical manual intervention and evaluation. Here we describe the results and performance of a novel interpretation software, which was used to characterize more than 50 recombinant serine/threonine kinases, receptor tyrosine kinases, or cytoplasmatic tyrosine kinases. LC-MS/MS data was acquired after tryptic digestion of the recombinantly produced kinases. The datasets have been imported to the proteome bioinformatics platform ProteinScape. The spectra were screened for a set of modifications, amino acid substitutions, unsuspected large measurement errors, enzyme no-specificity, and unknown mass shifts. The software restricts the search space by testing only sequences of interest. In widely used sequence database searches, testing all modifications and possible non-specific cleavages is not feasable. Besides the increase in sequence coverage basically caused by detection of one side non-specifically cleaved peptides, numerous modifications were found—namely, phosphorylation. methylation, pyroglutamate formation, methinonine oxidation, and N-terminal acetylation. As spectra of phosphorylated peptides are almost always in the minority compared to their unmodified counterparts, their detection is a challenge, but internal significance analysis revealed a substantial amount of phosphorylation.The phenomenon of auto-phosphorylation of kinase proteins was successfully monitored. The phosphorylation sites are categorized according to their sequence motive, and additionally their distribution is compared to phosphoryation sites described in public databases. Using this software triggered by the proteome database software proteinscape, searches were performed in a highly automated manner. Manual analysis could be reduced to minutes for the LC-MS/MS datasets containing more than 1000 spectra. Integrated result presentation strategies, which use clustering of spectra results on the amino acid level to annotate the protein sequence of interest, avoided the the possibility of seeing excess PTM contained in the large amount of acquired spectra.
The Bioinformatics Committee of the HUPO Brain Proteome Project (HUPO BPP) meets regularly to execute the post-lab analyses of the data produced in the HUPO BPP pilot studies. On January 9-11, 2006 the members as well as invited analysts came together at the European Bioinformatics Institute in Hinxton, UK for the pilot studies jamboree. The results of the reprocessing were presented and tasks forces were initiated to compile, to interpret and to summarise the data obtained.
In proteomics workflows, proteins are often digested first, then peptides are separated and subjected to identification by mass spectrometry (e.g., 2D-LC). In this process the peptide assignment to a protein is lost and has to be rebuilt by bioinformatic methods. We present ProteinExtractor, a module of the ProteinScape Bioinformatics Platform, which uses an empiric, iterative method to derive minimal protein lists from peptide search results, which may even come from different search algorithms or different MS datasets. ProteinExtractor uses an iterative approach to generate a minimal protein list. With composite database searches ProteinExtractor allows measuring the false-positive rate of the protein list. A test dataset (five recombinant proteins, 408 spectra, Bruker Ultraflex), and a real-life dataset (200410 LC/ESI-MS/MS spectra, Bruker Esquire HCT-Ultra, and 11619 LC/MALDI-MS/MS spectra, Bruker Ultraflex, both obtained from an analysis of proteins from a human cell line—SW480) were analyzed. The most probable protein sequence entries contained in the test dataset were identified with intensive manual data interpretation by several mass spectrometry experts. Using standard search algorithms, the correct protein sequence database entries are scattered over the first 171 protein ranks. Together with application specialists, we developed a set of rules to define a minimal protein list containing only those proteins (and isoforms) that can be unequivocally distinguished on the basis of MS/MS data. Applying these rules, the correct five proteins are ranked within the top eight protein candidates. In the real-life dataset, the peptide search results of Mascot, Sequest, Phenyx, and ProteinSolver were merged using ProteinExtractor. Merging all four search algorithms, over 50% more proteins could be identified than by using Mascot alone (with a false-positive rate of less then 2.5%). Merging ESI and MALDI data together, another 25% more proteins could be identified.
The newly available techniques for sensitive proteome analysis and the resulting amount of data require a new bioinformatics focus on automatic methods for spectrum reprocessing and peptide/protein validation. Manual validation of results in such studies is not feasible and objective enough for quality relevant interpretation. The necessity for tools enabling an automatic quality control is, therefore, important to produce reliable and comparable data in such big consortia as the Human Proteome Organization Brain Proteome Project. Standards and well-defined processing pipelines are important for these consortia. We show a way for choosing the right database model, through collecting data, processing these with a decoy database and end up with a quality controlled protein list merged from several search engines, including a known false-positive rate.
A novel software tool named PTM-Explorer has been applied to LC-MS/MS datasets acquired within the Human Proteome Organisation (HUPO) Brain Proteome Project (BPP). PTM-Explorer enables automatic identification of peptide MS/MS spectra that were not explained in typical sequence database searches. The main focus was detection of PTMs, but PTM-Explorer detects also unspecific peptide cleavage, mass measurement errors, experimental modifications, amino acid substitutions, transpeptidation products and unknown mass shifts. To avoid a combinatorial problem the search is restricted to a set of selected protein sequences, which stem from previous protein identifications using a common sequence database search. Prior to application to the HUPO BPP data, PTM-Explorer was evaluated on excellently manually characterized and evaluated LC-MS/MS data sets from Alpha-A-Crystallin gel spots obtained from mouse eye lens. Besides various PTMs including phosphorylation, a wealth of experimental modifications and unspecific cleavage products were successfully detected, completing the primary structure information of the measured proteins. Our results indicate that a large amount of MS/MS spectra that currently remain unidentified in standard database searches contain valuable information that can only be elucidated using suitable software tools.
The HUPO Brain Proteome Project is an initiative coordinating proteomics studies to characterise human and mouse brain proteomes. Proteins identified in human brain samples during the project's pilot phase were put into biological context through integration with various annotation sources followed by a bioinformatics analysis. The data set was related to the genome sequence via the genes encoding identified proteins including an assessment of splice variant identification as well as an analysis of tissue specificity of the respective transcripts. Proteins were furthermore categorised according to subcellular localisation, molecular function and biological process, grouped into protein families and mapped to biological pathways they are known to act in. Involvement in pathological conditions was examined based on association with entries in the online version of Mendelian Inheritance in Man and an interaction network was derived from curated protein‐proteininteraction data. Overall a non‐redundant set of 1804 proteins was identified in human brain samples. In the majority of cases splice variants could be unambiguously identified by unique peptides, including matches to several hypothetical transcripts of known as well as predicted genes.
Within the pilot phase of the HUPO Brain Proteome Project, nine participating laboratories analysed human (epilepsy and/or post mortem material) and mouse brain samples (embryonic, juvenile and adult), respectively, using a variety of different state of the art techniques. Thirty-seven different analytical approaches were accomplished. Of these analyses, 17 were done differentially, i.e. the protein expression patterns of the different samples (human or mouse) were compared. A catalogue of all proteins present in the respective sample was built in 20 analyses (mapping). All data were collected in the Data Collection Center in Bochum, Germany, and were reprocessed according to thoroughly defined parameters. In this report, a summary of all results and inter-laboratory comparisons with respect to the number of identified proteins, the analysed organism, and the used techniques is presented.
In the present work the complexity in the 2D-gel protein pattern of murin lenticular αA-Crystallin was analyzed. An in depth study of the different protein isoforms was done combining different proteomic tools. Lens proteins of four different ages, from embryo to 100-week-old mice, were separated by large 2D-PAGE, revealing an increase in the number and intensity of the spots of αA-Crystallin during the process of aging. For further analyses the oldest mice were chosen. Comparison and evaluation of two different staining methods proved Imidazole-Zinc to be a good alternative to the generally used Coomassie stain. The characterization of the different αA-Crystallin protein species was done using nanoLC-ESI-MS/MS (liquid chromatography electrospray ionisation tandem mass spectrometry). Data interpretation was done by database searching, manual validation and a new MS/MS-interpretation tool for posttranslational modifications—the PTM-Explorer. Using this way, eight different phosphorylation sites were identified and localized; the identification of four of them was not published so far. Furthermore, quantitative N-terminal acetylation of αA-Crystallin and variable C-terminal truncation was observed, also not published in this extent yet. The results of the mass spectrometric analysis were validated by immunoblotting experiments using two different αA-Crystallin specific antibodies. In addition, a fluorescent phospho-specific stain was used to detect the protein spots including phosphorylation groups. Re-separation 2D-PAGE was done to round off the present study and explain the appearance of some of the protein spots in the gel as artifacts of the 2D-PAGE separation.