A method for constructing one-dimensional proteomic maps (1D-PM) based on mass spectrometric identification of proteins from adjacent slices of one-dimensional electrophoregram has been developed. For the proteomic mapping, gel lanes were sectioned into slices less than 0.2 mm thick and each slice was subjected to enzymatic hydrolysis. The resultant mixture of peptide fragments was analyzed by matrix-assisted laser desorption time-of-flight mass spectrometry (MALDI-TOF) and liquid chromatography electrospray ionization tandem mass spectrometry (LC-MS/MS). Proteins were identified by the mass spectra obtained. Data on peptide fragments and corresponding identified proteins were presented as a 1D-PM. Proteomic maps were constructed by assigning individual proteins to gel slices based on number of matching peptides in a corresponding MS-data. On 1D-PM of human liver microsomal fraction, 18 proteins were identified in the region of 40–65 kDa. These included 12 membrane proteins belonging to the superfamily of cytochromes P450. Pooling of mass spectrometric data, obtained from several adjacent gel slices (molecular zooming) increased sequence coverage of CYP2A (cytochrome P450 family 2A). The maximal coverage of 66% significantly exceeded the level of 48% that could be obtained using one (even the most informative) slice. This method can be applied to the proteomic profiling of membrane-bound proteins.
In proteomics workflows, proteins are often digested first, then peptides are separated and subjected to identification by mass spectrometry (e.g., 2D-LC). In this process the peptide assignment to a protein is lost and has to be rebuilt by bioinformatic methods. We present ProteinExtractor, a module of the ProteinScape Bioinformatics Platform, which uses an empiric, iterative method to derive minimal protein lists from peptide search results, which may even come from different search algorithms or different MS datasets. ProteinExtractor uses an iterative approach to generate a minimal protein list. With composite database searches ProteinExtractor allows measuring the false-positive rate of the protein list. A test dataset (five recombinant proteins, 408 spectra, Bruker Ultraflex), and a real-life dataset (200410 LC/ESI-MS/MS spectra, Bruker Esquire HCT-Ultra, and 11619 LC/MALDI-MS/MS spectra, Bruker Ultraflex, both obtained from an analysis of proteins from a human cell line—SW480) were analyzed. The most probable protein sequence entries contained in the test dataset were identified with intensive manual data interpretation by several mass spectrometry experts. Using standard search algorithms, the correct protein sequence database entries are scattered over the first 171 protein ranks. Together with application specialists, we developed a set of rules to define a minimal protein list containing only those proteins (and isoforms) that can be unequivocally distinguished on the basis of MS/MS data. Applying these rules, the correct five proteins are ranked within the top eight protein candidates. In the real-life dataset, the peptide search results of Mascot, Sequest, Phenyx, and ProteinSolver were merged using ProteinExtractor. Merging all four search algorithms, over 50% more proteins could be identified than by using Mascot alone (with a false-positive rate of less then 2.5%). Merging ESI and MALDI data together, another 25% more proteins could be identified.
The newly available techniques for sensitive proteome analysis and the resulting amount of data require a new bioinformatics focus on automatic methods for spectrum reprocessing and peptide/protein validation. Manual validation of results in such studies is not feasible and objective enough for quality relevant interpretation. The necessity for tools enabling an automatic quality control is, therefore, important to produce reliable and comparable data in such big consortia as the Human Proteome Organization Brain Proteome Project. Standards and well-defined processing pipelines are important for these consortia. We show a way for choosing the right database model, through collecting data, processing these with a decoy database and end up with a quality controlled protein list merged from several search engines, including a known false-positive rate.
Proteomics of membrane-bound proteins requires new technological approaches powered by the appropriate bioinformatic tools. Off-line molecular scanning is one of such approaches, that fully exploits the capacities provided by modern instruments and robotized lines. The SDS-PAGE separated liver microsomes are scanned by picking up the 1.5 mm spots closely one after another. Each picked spot is digested and peptide mass fingerprint is collected. The series of fingerprints is then processed to assemble the consistent 1D proteomic map.
The pilot phase of the Human Brain Proteome Project as a part of the Human Proteome Organisation has just been started. In two pilot studies, 18 different laboratories are analyzing mouse brains of three age stages and human brain autopsy versus biopsy material, respectively. The overall aim is to elucidate the portfolio of available techniques as well as to elaborate common standards. As a first step, it was decided to use the common bioinformatics platform ProteinScape™ that was introduced to the participating groups in a two day course in Bochum, Germany.