Mass spectrometers are tools in the field of proteomics. Each mass spectrometer manufacturer uses its individual data format and software tools, making the creation of additional software tools and databases difficult and incompatible with one another. The Mass Spectrum I/O Project (MSIOP) addresses this problem, allowing for storage and analysis of mass spectrometer data from multiple manufacturers across various platforms while providing a framework upon which mass spectrometry software tools can be constructed.
alpha-Synuclein (alpha-syn) is an intrinsically unstructured 140-residue neuronal protein of uncertain function that is implicated in the etiology of Parkinson's disease. Tertiary contact formation rate constants in (alpha-syn, determined from diffusion-limited electron-transfer kinetics measurements, are poorly approximated by simple random polymer theory. One source of the discrepancy between theory and experiment may be that interior-loop formation rates are not well approximated by end-to-end contact dynamics models. We have addressed this issue with Monte Carlo simulations to model asynchronous and synchronous motion of contacting sites in a random polymer. These simulations suggest that a dynamical drag effect may slow interior-loop formation rates by about a factor of 2 in comparison to end-to-end loops of comparable size. The additional deviations from random coil behavior in alpha-syn likely arise from clustering of hydrophobic residues in the disordered polypeptide.
A significant consequence of protein phosphorylation is to alter protein-protein interactions, leading to dynamic regulation of the components of protein complexes that direct many core biological processes. Recent proteomic studies have populated databases with extensive compilations of cellular phosphoproteins and phosphorylation sites and a similarly deep coverage of the subunit compositions and interactions in multiprotein complexes. However, considerably less data are available on the dynamics of phosphorylation, composition of multiprotein complexes or that define their interdependence. We describe a method to identify candidate phosphoprotein complexes by combining phosphoprotein affinity chromatography, separation by size, denaturing gel electrophoresis, protein identification by tandem mass spectrometry, and informatics analysis. Toward developing phosphoproteome profiling, we have isolated native phosphoproteins using a phosphoprotein affinity matrix, Pro-Q Diamond resin (Molecular Probes-Invitrogen). This resin quantitatively retains phosphoproteins and associated proteins from cell extracts. Pro-Q Diamond purification of a yeast whole cell extract followed by 1-D PAGE separation, proteolysis and ESI LC-MS/MS, a method we term PA-GeLC-MS/MS, yielded 108 proteins, a majority of which were known phosphoproteins. To identify proteins that were purified as parts of phosphoprotein complexes, the Pro-Q eluate was separated into two fractions by size, <100 kDa and >100 kDa, before analysis by PAGE and ESI LC-MS/MS and the component proteins queried against databases to identify protein-protein interactions. The <100 kDa fraction was enriched in phosphoproteins indicating the presence of monomeric phosphoproteins. The >100 kDa fraction contained 171 proteins of 20-80 kDa, nearly all of which participate in known protein-protein interactions. Of these 171, few are known phosphoproteins, consistent with their purification by participation in protein complexes. By comparing the results of our phosphoprotein profiling with the informational databases on phosphoproteomics, protein-protein interactions and protein complexes, we have developed an approach to examining the correlation between protein interactions and protein phosphorylation.
Despite advances in methods and instrumentation for analysis of phosphopeptides using mass spectrometry, it is still difficult to quantify the extent of phosphorylation of a substrate because of physiochemical differences between unphosphorylated and phosphorylated peptides. Here we report experiments to investigate those differences using MALDI-TOF mass spectrometry for a set of synthetic peptides by creating calibration curves of known input ratios of peptides/phosphopeptides and analyzing their resulting signal intensity ratios. These calibration curves reveal subtleties in sequence-dependent differences for relative desorption/ionization efficiencies that cannot be seen from single-point calibrations. We found that the behaviors were reproducible with a variability of 5-10% for observed phosphopeptide signal. Although these data allow us to begin addressing the issues related to modeling these properties and predicting relative signal strengths for other peptide sequences, it is clear that this behavior is highly complex and needs to be further explored.
B43 Molecularly targeted cancer therapies such as antibodies and small molecule signal transduction inhibitors have proved extraordinary value both as single agents and adjuncts to conventional cytotoxic therapy. As examples, hormonal blockade slows progression of breast and prostate cancer while kinase inhibitors have found wide use in treating leukemias and solid tumors. Yet, few agents demonstrate the efficacy and/or safety to reach Phase III trials or clinical use, reflecting off-target effects, patient-to-patient variability and other factors. Clearly, better assays for diagnosis, prediction and monitoring are needed. While simple biomarkers have proved inadequate, molecular signatures that predict tumor sensitivity and report on efficacy or resistance remain to be developed. Multiplexed analysis of gene expression is robust, but targets may be far upstream of transcription. A proteomic approach focused on analysis of signal transduction is an attractive alternative. We have recently developed a straightforward method for definition of global phosphoprotein profiles that may help answer the challenge. Our methodology, dubbed PA-GeLC-MS/MS, involves phosphoprotein affinity purification, separation by SDS-PAGE, LC-MS/MS and informatic analysis. We exploit stable isotope labeling to differentiate phosphoprotein abundance in pairs of samples, allowing direct comparison of controls with cells treated with targeted agents. We have implemented novel informatic tools to exploit high mass accuracy and isotopic labels that enhance speed and reproducibility of protein identification and quantitation. To validate our approach, we performed phosphoprotein proteome profiling of K562 Ph1+ chronic myelogenous leukemia cells before and after treatment with the Bcr-Abl tyrosine kinase inhibitor imatinib. Phosphoprotein affinity purification led to enrichment of both tyrosine and serine/threonine phosphorylated proteins. LC-MS/MS analysis identified nearly 2000 proteins from a single gel lane with high confidence. Abundance analysis revealed many proteins differentially regulated by imatinib, including several known Bcr-Abl substrates. Toward applying this approach to circulating CML cells, we will compare imatinib with the second generation agents dasatinib and nilotinib and develop signatures for imatinib resistance. Following a similar path, to reveal markers of tamoxifen sensitivity and resistance, we are profiling breast cancer cell lines based on phosphoprotein enrichment of cytoplasmic proteins. We intend to define signatures of cellular signaling that distinguish malignancies and report effects of targeted agents, with the ultimate goal of developing high-throughput assays compatible with drug discovery and clinical testing.
With the increasing amount of work being performed in the field of Mass Spectrometry (MS), a huge amount of data is being generated. This data needs to be properly managed, organized and shared among researchers at various institutions. The problem is further complicated by the different proprietary formats used by manufacturers of MS machines. We demonstrate an end-to- end system to automate the process of converting the data to an open format, and to upload the data to a centralized server where it is easily organized and managed. The system allows scientists to browse, download, and use the data with third party tools. The user-view is simple and hides the underlying data-management system.
Managing the I/O and organization of mass spectrometry data in the form of files and structures within a program is often time consuming due to the many fields that need to be set and retrieved. Each programmer handling this in his/her own way leads to incompatibilities and inefficiencies between implementations. The Mass Spectrometry I/O Project addresses these issues by creating a framework that handles mass spectrometry data I/O and data organization thereby allowing researchers to concentrate on data analysis rather than I/O. In addition, the Mass Spectrometry I/O Project leverages several cross platform and portability enhancing technologies, allowing it to be utilized on a wide variety of hardware and operating system combinations.
Workflow management is an important part of scientific experiments. A common pattern that scientists are using is based on repetitive job execution on a variety of different systems, and managing such job execution is necessary for large-scale scientific workflows. The workflow system should also be client-based and able to handle multiple security contexts to allow researchers to take advantage of a diverse array of systems. We have developed, based on the Java Commodity Grid Kit (CoG Kit), a sophisticated and extensible workflow system that seamlessly integrates with the queueing system Cobalt through the advanced features provided by the CoG Kit.
Mass spectrometry (MS) is an analytical method biologists and chemists use to sequence and identify unknown proteins. Although sensitivity and sample analysis modalities differ, mass spectrometers basically consist of an ion source, a mass analyzer, and a detector. The final output is a single mass spectrum or a series of mass spectra depending on the equipment used. Each spectrum consists of pairs of mass-to-charge ratio (m/z) and intensity [1]. In the process of capturing data for each spectrum, noise from the spectrometer appears as m/z and intensity pairs interspersed with those from the analyzed sample (analyte). Electronic noise can come from sensors, light sources, and other internal electrical systems. Chemical noise specifically refers to ions that do not belong to the analyte. Chemical noise can come from analyte preparation or contamination. Noise can have a detrimental effect on the performance of algorithms that assist protein identification such as de novo sequencing, spectral scoring used in comparisons, and database searching. Effective reduction or removal of such noise is critical for improved results. We propose refinement of current wavelet-based techniques to improve noise removal in MS data. These wavelet-based techniques will be incorporated into a workflow being developed on the Illinois Bio-Grid (IBG) [2]. Since noise appears as seemingly ordinary data, removing it can be very difficult. Noise removal based on simple global thresholds would remove noise and less intense peaks, losing potentially useful signals. Multiscale analysis methods such as wavelets have shown promise, removing low intensity noise while preserving medium and higher intensity peaks [3]. The standard procedure for denoising with wavelets starts by performing the transform on the raw data. The wavelet transform decomposes the data into a series of average and detail values which are used during reconstruction of the data. The transformed data is thresholded by attenuating signals below a user-defined noise level. Determining a good threshold value will require experimentation with known samples. An inverse transform is done on the thresholded transform data to reconstruct a good approximation of the denoised data. Measuring the quality of the denoising process will require repeated testing with objective quality measures such as RMS error and signal-to-noise ratio measurements. The guidance of an experienced chemist helps the refinement of the denoising algorithms. Since wavelet transforms can preserve higher intensities very well, significant lower intensity peaks will be watched closely. Lower intensity peaks from raw data, especially those along the baseline, can be unintentionally eliminated during denoising and require further investigation [4]. Using wavelet-based denoising is part of a larger workflow that includes baseline correction, peak detection, and integration. Ongoing research has produced a variety of algorithms for each step in the workflow. The ability to choose different algorithms in the workflow affords the developer, researcher, or technician the chance to experiment and choose the workflow that best suits their data. For larger scale deployments and high-throughput analysis, an established workflow improves consistency. Baseline correction fits MS data to a straight line to ease the integration (area under the curve) of peaks during the quantification process. Williams, et al. correct using a polynomial fitted along medians within windows on the m/z axis [5], while Liu, et al. correct by subtracting the convex hull of the spectrum after smoothing raw data [6]. Peak detection is critical to proper protein identification and can depend on the quality of the denoised signal. It is complicated by false and merged peaks, which can be deconvoluted using peak width constraints or wavelet decomposition.
article Challenges in systems biology: an interview with Dr. Lee Hood, inventor of the DNA sequencer Share on Author: David Sigfredo Angulo View Profile Authors Info & Claims XRDS: Crossroads, The ACM Magazine for StudentsVolume 13Issue 1September 2006 pp 8https://doi.org/10.1145/1217666.1217674Online:01 September 2006Publication History 0citation153DownloadsMetricsTotal Citations0Total Downloads153Last 12 Months0Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
In this paper, we address the problem of searching huge biological databases on the scale of at least several gigabytes by utilizing parallel processing. Biological databases storing DNA sequences, protein sequences, or mass spectra are growing exponentially. Searches through these databases consume exponentially growing computational resources as well. We demonstrate herein a general use, MPI based, C++ framework for generically splitting databases amongst several computational nodes. The combined RAM of the nodes working in tandem is often sufficient to keep the entire database in memory, and therefore to search it efficiently without paging to disk. The framework runs as a persistent service, processing all submitted queries. This allows for query reordering and better utilization of the memory. Thereby, we achieve superlinear speedups compared to single processor implementations. We demonstrate the utility and speedup of the framework using a real biological database and an actual searching algorithm for mass spectrometry.
The Inverse Protein Folding problem, also called Protein Design, spans the boundaries of both the computer and biological sciences. The problem consists of determining a sequence of amino acids to compose a protein that will, due to their combined bio-chemical properties, fold into a predetermined three-dimensional structure. This three-dimensional structure, or conformation, determines the effect of the protein on its environment. Hence, success in the Inverse Protein Folding problem would allow scientists the ability to redesign existing proteins with additional or enhanced functionality or even to design new proteins with novel functionality. Nevertheless, while the impact of a solution to the Inverse Protein Folding is apparent, the computational challenges are extraordinary, since the search space explodes exponentially as the length of the protein increases. We demonstrate a Monte-Carlo model for sampling the search space coupled with energy functions for evaluating conformations. These energy functions are based on statistical propensities mined from the Protein Data Bank. We further demonstrate a framework based on the Illinois Bio-Grid Toolkit which utilizes grid technologies to give the capabilities of engaging massive amounts of computational power to the problem.
In this paper, we introduce a framework for experi- ment management that simplifies the user's interac- tion with Grid environments. We have developed a service that allows the individual scientist to manage a large number of tasks as typically found in exper- iment management. Our service includes the ability to conduct application state notifications. Similar to the definition of standard output and standard er- ror, we have defined standard status that allows us to conduct application status notifications. We have tested our tool with a large number of long running experiments, and shown its usability in practical ap- plications such as bioinformatics.
Research in proteomics has created two significant needs: the need for an accurate public database of empirically derived mass spectrum information and the need for managing the I/O and organization of mass spectrometry data in the form of files and structures. Lack of an empirically derived database limits the ability of proteomic researchers to identify and study proteins. Managing the I/O and organization of mass spectrometry data is often time-consuming due to the many fields that need to be set and retrieved. As a result, incompatibilities and inefficiencies are created by each programmer handling this in his or her own way. Until recently, storage space and computing power has been the limiting factor in developing tools to handle the vast amount of mass spectrometry information. Now the resources are available to store, organize, and analyze mass spectrometry information.The Illinois Bio-Grid Mass Spectrometry Database is a database of empirically derived tandem mass spectra of peptides created to provide researchers with an organized and searchable database of curated spectrum information to allow more accurate protein identification. The Mass Spectrometry I/O Project creates a framework that handles mass spectrometry data I/O and data organization, allowing researchers to concentrate on data analysis rather than I/O. In addition, the Mass Spectrometry I/O Project leverages several cross-platform and portability-enhancing technologies, allowing it to be utilized on a variety of hardware and operating systems.
Proteomic researchers who study mass spectrometry data have expressed a need for an accurate public database of empirically derived curated mass spectrum information. Lack of such a database limits proteomic researchers in their ability to identify and study proteins. Until recently, storage space and computing power has been the limiting factor in developing tools to handle the vast amount of mass spectrometry information. Now, the resources are available to store, organize, and analyze mass spectrometry information. The Illinois BioGrid Mass Spectrometry Database is a database of empirically derived tandem mass spectra of peptides created to provide researchers with an organized and searchable database of curated spectrum information to allow more accurate protein identification. This paper will discuss the methods used to import the mass spectra into the Illinois Bio-Grid Mass Spectrometry Database, as well as the database requirements, motivation, use cases, design, and results.
Database searching for protein identification is an efficient approach for MS/MS spectra identification compared to the existing approaches, such as protein identification with antibodies, and chemical degradation. In order for a database search to give reliable results, the database itself, as well as the searching algorithm, must be reliable. However, MS/MS spectra identification is still imperfect. The Illinois Bio-Grid Mass Spectrometry Database (IBG-MSD), an empirically derived curated and annotated database, along with multiple searching algorithms, addresses this issue and provides a better identification tool. The currently implemented algorithms are K-mutation algorithm, spectral contrast angle, and similarity index algorithm. The search engine allows users to search for post-translationally modified peptide using the K-mutation algorithm implemented. We provide a framework for high speed query processing and for efficient identification task. This new framework allows a community of users integratively to build a larger database and provide more efficient protein identification tools.
The manipulation of codon selection in synthetic genes greatly influences expression of the gene's products i.e. proteins. Biologists frequently have multiple objectives when constructing a synthetic gene for a given protein sequence and are often limited by the tools and mechanisms used to design genes in vitro. Poor codon choice and mRNA secondary structure formation can lead to to the under-expression of a protein coded for by a synthetic gene; oligonucleotide melting temperature and the existence of restriction sites can either facilitate or inhibit gene assembly processes. We have attempted to capture these heuristics in a genetic algorithm that generates optimized DNA sequences for a given input amino acid sequence. Our preliminary results indicate that the algorithm is very effective at quickly creating genes which match desired DNA sequence characteristics.
This month, Crossroads brings you a special issue devoted to bioinformatics the computational analysis of biological data. Included in this issue is an interview with one of the field's foremost visionaries, Dr. Lee Hood of the Institute of Systems Biology and inventor of the DNA sequencer. Dr. Hood discusses the history of bioinformatics and its recent challenges, including the difficulties of funding in the current political climate. We also present four articles covering recent research:
While distributed, heterogeneous collections of computers ("Grids") can in principle be used as a computing platform, in practice the problems of first discovering and then organizing resources to meet application requirements are difficult. We present a general-purpose resource selection framework that addresses these problems by defining a resource selection service for locating Grid resources that match application requirements. At the heart of this framework is a simple, but powerful, declarative language based on a technique called set matching, which extends the Condor matchmaking framework to support both single-resource and multiple-resource selection. This framework also provides an open interface for loading application-specific mapping modules to personalize the resource selector. We present results obtained when this framework is applied in the context of a computational astrophysics application, Cactus. These results demonstrate the effectiveness of our technique.
The ability to harness heterogeneous, dynamically available grid resources is attractive to typically resource-starved computational scientists and engineers, as in principle it can increase, by significant factors, the number of cycles that can be delivered to applications. However, new adaptive application structures and dynamic runtime system mechanisms are required if we are to operate effectively in grid environments. To explore some of these issues in a practical setting, the authors are developing an experimental framework, called Cactus, that incorporates both adaptive application structures for dealing with changing resource characteristics and adaptive resource selection mechanisms that allow applications to change their resource allocations (e.g., via migration) when performance falls outside specified limits. The authors describe the adaptive resource selection mechanisms and describe how they are used to achieve automatic application migration to “better” resources following performance degradation. The results provide insights into the architectural structures required to support adaptive resource selection. In addition, the authors suggest that the Cactus Worm affords many opportunities for grid computing.
G. Von Laszewski合作论文数Indiana University2