This paper describes the methodology of providing multiprobability predictions for proteomic mass spectrometry data. The methodology is based on a newly developed machine learning framework called Venn machines. Is allows to output a valid probability interval. The methodology is designed for mass spectrometry data. For demonstrative purposes, we applied this methodology to MALDI-TOF data sets in order to predict the diagnosis of heart disease and early diagnoses of ovarian cancer and breast cancer. The experiments showed that probability intervals are narrow, that is, the output of the multiprobability predictor is similar to a single probability distribution. In addition, probability intervals produced for heart disease and ovarian cancer data were more accurate than the output of corresponding probability predictor. When Venn machines were forced to make point predictions, the accuracy of such predictions is for the most data better than the accuracy of the underlying algorithm that outputs single probability distribution of a label. Application of this methodology to MALDI-TOF data sets empirically demonstrates the validity. The accuracy of the proposed method on ovarian cancer data rises from 66.7 % 11 months in advance of the moment of diagnosis to up to 90.2 % at the moment of diagnosis. The same approach has been applied to heart disease data without time dependency, although the achieved accuracy was not as high (up to 69.9 %). The methodology allowed us to confirm mass spectrometry peaks previously identified as carrying statistically significant information for discrimination between controls and cases.
This paper describes the methodology of providing multiprobability predictions for proteomic mass spectrometry data. The methodology is based on a newly developed machine learning framework called Venn machines. They allow us to output a valid probability interval. We apply this methodology to mass spectrometry data sets in order to predict the diagnosis of heart disease and early diagnoses of ovarian cancer. The experiments show that probability intervals are valid and narrow. In addition, probability intervals were compared with the output of a corresponding probability predictor.
The work describes an application of a recently developed machine-learning technique called Mondrian predictors to risk assessment of ovarian and breast cancers. The analysis is based on mass spectrometry profiling of human serum samples that were collected in the United Kingdom Collaborative Trial of Ovarian Cancer Screening. The work describes the technique and presents the results of classification (diagnosis) and the corresponding measures of confidence of the diagnostics. The main advantage of this approach is a proven validity of prediction. The work also describes an approach to improve early diagnosis of ovarian and breast cancers since the data in the United Kingdom Collaborative Trial of Ovarian Cancer Screening were collected over a period of 7 years and do allow to make observations of changes in human serum over that period of time. Significance of improvement is confirmed statistically (for up to 11 months for ovarian cancer and 9 months for breast cancer). In addition, the methodology allowed us to pinpoint the same mass spectrometry peaks as previously detected as carrying statistically significant information for discrimination between healthy and diseased patients. The results are discussed.
The paper describes an application of a recently developed machine learning technique called Mondrian predictors to risk assessment of ovarian and breast cancers. The analysis is based on mass spectrometry profiling of human serum samples that were collected in the United Kingdom Collaborative Trial of Ovarian Cancer Screening. The paper describes the technique and presents the results of classification (diagnosis) and the corresponding measures of confidence of the diagnostics. The main advantage of this approach is a proven validity of prediction. The paper also describes an approach to improve early diagnosis of ovarian and breast cancers since the data in the United Kingdom Collaborative Trial of Ovarian Cancer Screening were collected over a period of seven years and do allow to make observations of changes in human serum over that period of time. Significance of improvement is confirmed statistically (for up to 11 months for Ovarian Cancer and 9 months for Breast Cancer). In addition, the methodology allowed us to pinpoint the same mass spectrometry peaks as previously detected as carrying statistically significant information for discrimination between healthy and diseased patients. The results are discussed. 1 Computer Learning Research Centre, Royal Holloway, University of London 2 EGA Institute for Women’s Health, University College London 3 BioCentre and Department of Chemistry, University of Reading ∗ Address correspondence to Dr Ilia Nouretdinov, Department of Computer Science, Royal Holloway, University of London, Egham, Surrey TW20 0EX. Tel/Fax: +44(0)1784 443912/439786. Email: ilia@cs.rhul.ac.uk.
Knowledge of the differences between the amounts and types of protein that are expressed in diseased compared to healthy subjects may give an understanding of the biological pathways that cause disease. This is the reasoning behind the presented protocol, which uses difference gel electrophoresis (DIGE) to discover up- or down-regulated proteins between mice of different genotypes, or of those fed on different diets, that may thus be prone to develop diabetes-like phenotypes. Subsequent analysis of these proteins by tandem mass spectrometry typically facilitates their identification with a high degree of confidence.
Aim: A nested case-control discovery study was undertaken 10 test whether information within the serum peptidome can improve on the utility of CA125 for early ovarian cancer detection. Materials and Methods: High-throughput matrix-assisted laser desorption ionisation mass spectrometry (MALDI-MS) was used to profile 295 serum samples from women pre-dating their ovarian cancer diagnosis and from 585 matched control samples. Classification rules incorporating CA125 and MS peak intensities were tested for discriminating ability. Results: Two peaks were found which in combination with CA125 discriminated cases from controls up to 15 and 11 months before diagnosis, respectively, and earlier than using CA125 alone. One peak was identified as connective tissue-activating peptide III (CTAPIII), whilst the other was putatively identified as platelet factor 4 (PF4). ELISA data supported the down-regulation of PF4 in early cancer cases. Conclusion: Serum peptide information with CA125 improves lead time for early detection of ovarian cancer. The candidate markers are platelet-derived chemokines, suggesting a link between platelet function and tumour development.
BACKGROUND:Diabetes like many diseases and biological processes is not mono-causal. On the one hand multi-factorial studies with complex experimental design are required for its comprehensive analysis. On the other hand, the data from these studies often include a substantial amount of redundancy such as proteins that are typically represented by a multitude of peptides. Coping simultaneously with both complexities (experimental and technological) makes data analysis a challenge for Bioinformatics.RESULTS:We present a comprehensive work-flow tailored for analyzing complex data including data from multi-factorial studies. The developed approach aims at revealing effects caused by a distinct combination of experimental factors, in our case genotype and diet. Applying the developed work-flow to the analysis of an established polygenic mouse model for diet-induced type 2 diabetes, we found peptides with significant fold changes exclusively for the combination of a particular strain and diet. Exploitation of redundancy enables the visualization of peptide correlation and provides a natural way of feature selection for classification and prediction. Classification based on the features selected using our approach performs similar to classifications based on more complex feature selection methods.CONCLUSIONS:The combination of ANOVA and redundancy exploitation allows for identification of biomarker candidates in multi-dimensional MALDI-TOF MS profiling studies with complex experimental design. With respect to feature selection our method provides a fast and intuitive alternative to global optimization strategies with comparable performance. The method is implemented in R and the scripts are available by contacting the corresponding author.
Differential MS analysis of blood samples from diseased and control subjects is increasingly being employed in the hunt for biomarkers that can detect disease at an early stage. For diagnostic tests, in particular for population screening, robust protocols are required that can offer high-throughput analysis, ideally at high mass spectrometric sensitivity. To achieve this, blood samples need to be collected, prepared and analyzed in a standardized manner that minimizes potential bias. Simple purification methods combined with MALDI MS profiling have so far been championed for providing the best approach. In this chapter, we describe an adapted and validated protocol based on a simple and fast solid-phase extraction technique using ZipTips (R). This protocol facilitates the purification of potential blood biomarkers in a few steps for mass spectral biomarker pattern diagnostics using MALDI. It is suitable for use in an automated high-throughput and potentially clinical environment and has the advantage of only requiring a few microlitres of blood plasma or serum. The presented protocol has been tested over several years in our laboratory and found to be more reproducible and suitable for plasma and serum profiling than similar methodologies based on magnetic bead purification.
Hydroponic cultivation is a soil-free technique commonly used in agriculture where plants are grown in a highly controlled environment leading to high crop yields. Hydroponic Isotope Labeling of Entire Plants (HILEP) is the most cost-effective isotope labeling method for quantitative plant proteomics, enabling the metabolic labeling of whole and mature plants with a stable isotope such as N-15. Employing inorganic N-15-containing salts as the sole nitrogen source, healthy plants can be easily grown and labeled in hydroponic solutions. Close to 100% N-15-labeling of proteins can be achieved using HILEP. Moreover, hydroponic cultivation allows tight control of growth conditions. Plants grown in N-14- and N-15-hydroponic media are typically pooled straight after harvest, eliminating any bias due to subsequent sample preparation and analysis. The pooled N-14-/N-15-protein extracts can be fractionated in any convenient way and digested with trypsin (or any other enzyme of choice). Peptides can then be analyzed by techniques such as liquid chromatography electrospray ionization tandem mass spectrometry (LC-ESI-MS/MS). Following protein identification, the spectra of N-14/N-15-peptide pairs are typically compared and relative protein amounts are calculated from the N-14/N-15-ion signal ratios. An increasing number of bioinformatics tools are now available for determining these ratios in a convenient way.
Use of superdihydroxybenzoic acid as the matrix enabled the analysis of highly complex mixtures of proanthocyanidins from sainfoin (Onobrychis viciifolia) by MALDI-TOF mass spectrometry. Proanthocyanidins contained predominantly B-type homopolymers and heteropolymers up to 12-mers (3400 Da). Use of another matrix, 2,6-dihydroxyacetophenone, revealed the presence of A-type glycosylated dimers. In addition, we report here how a comparison of the isotopic adduct patterns, which resulted from Li and Na salts as MALDI matrix additives, could be used to confirm the presence of A-type linkages in complex proanthocyanidin mixtures. Preliminary evidence suggested the presence of A-type dimers in glycosylated prodelphinidins and in tetrameric procyanidins and prodelphinidins.
MALDI MS profiling, using easily available body fluids such as blood serum, has attracted considerable interest for its potential in clinical applications. Despite the numerous reports on MALDI MS profiling of human serum, there is only scarce information on the identity of the species making up these profiles, particularly in the mass range of larger peptides. Here, we provide a list of more than 90 entries of MALDI MS profile peak identities up to 10 kDa obtained from human blood serum. Various modifications such as phosphorylation were detected among the peptide identifications. The overlap with the few other MALDI MS peak lists published so far was found to be limited and hence our list significantly extends the number of identified peaks commonly found in MALDI MS profiling of human blood serum.
Objectives: Our objective was to test the performance of CA125 in classifying serum samples from a cohort of malignant and benign ovarian cancers and age-matched healthy controls and to assess whether combining information from matrix-assisted laser desorption/ionization (MALDI) time-of-flight profiling could improve diagnostic performance.Materials and Methods: Serum samples from women with ovarian neoplasms and healthy volunteers were subjected to CA125 assay and MALDI time-of-flight mass spectrometry (MS) profiling. Models were built from training data sets using discriminatory MALDI MS peaks in combination with CA125 values and tested their ability to classify blinded test samples. These were compared with models using CA125 threshold levels from 193 patients with ovarian cancer, 290 with benign neoplasm, and 2236 postmenopausal healthy controls.Results: Using a CA125 cutoff of 30 U/mL, an overall sensitivity of 94.8% (96.6% specificity) was obtained when comparing malignancies versus healthy postmenopausal controls, whereas a cutoff of 65 U/mL provided a sensitivity of 83.9% (99.6% specificity). High classification accuracies were obtained for early-stage cancers (93.5% sensitivity). Reasons for high accuracies include recruitment bias, restriction to postmenopausal women, and inclusion of only primary invasive epithelial ovarian cancer cases. The combination of MS profiling information with CA125 did not significantly improve the specificity/accuracy compared with classifications on the basis of CA125 alone.Conclusions: We report unexpectedly good performance of serum CA125 using threshold classification in discriminating healthy controls and women with benign masses from those with invasive ovarian cancer. This highlights the dependence of diagnostic tests on the characteristics of the study population and the crucial need for authors to provide sufficient relevant details to allow comparison. Our study also shows that MS profiling information adds little to diagnostic accuracy. This finding is in contrast with other reports and shows the limitations of serum MS profiling for biomarker discovery and as a diagnostic tool.
Background: The serum peptidome may be a valuable source of diagnostic cancer biomarkers. Previous mass spectrometry (MS) studies have suggested that groups of related peptides discriminatory for different cancer types are generated ex vivo from abundant serum proteins by tumor-specific exopeptidases. We tested 2 complementary serum profiling strategies to see if similar peptides could be found that discriminate ovarian cancer from benign cases and healthy controls. Methods: We subjected identically collected and processed serum samples from healthy volunteers and patients to automated polypeptide extraction on octadecylsilane-coated magnetic beads and separately on ZipTips before MALDI-TOF MS profiling at 2 centers. The 2 platforms were compared and case control profiling data analyzed to find altered MS peak intensities. We tested models built from training datasets for both methods for their ability to classify a blinded test set. Results: Both profiling platforms had CVs of approximately 15% and could be applied for high-throughput analysis of clinical samples. The 2 methods generated overlapping peptide profiles, with some differences in peak intensity in different mass regions. In cross-validation, models from training data gave diagnostic accuracies up to 87% for discriminating malignant ovarian cancer from healthy controls and up to 81% for discriminating malignant from benign samples. Diagnostic accuracies up to 71% (malignant vs healthy) and up to 65% (malignant vs benign) were obtained when the models were validated on the blinded test set. Conclusions: For ovarian cancer, altered MALDI-TOF MS peptide profiles alone cannot be used for accurate diagnoses.
Blood represents a convenient diagnostic specimen for clinical analysis. Serum metabolites and peptides have the potential to serve as reliable indicators of progression from a normal to a diseased state. However, the presence of salts, lipids and high concentrations of protein in blood serum can adversely affect its mass spectrometric profiling, as all of these moieties can hinder the ionisation and detection of diagnostic biomarkers. Several pre-fractionation strategies using chromatographic adsorbents have been employed to desalt samples and remove abundant proteins such as albumin and immunoglobulin. As an alternative to multi-stage fractionation chromatography, we describe in this practical review the adaptation and validation of a simple and fast solid-phase extraction technique using ZipTips. The protocol allows for the purification and concentration of peptides in a few, mostly automated steps, prior to spectral biomarker pattern diagnostics using MALDI MS and MS/MS. It has the added advantages of being suitable for use in a high-throughput and potentially clinical environment, only requiring a few microlitres of serum. We have evaluated, optimised, and standardised a number of analytical parameters ranging from serum storage and handling to automated peptide extraction, crystallisation, spectral acquisition, and signal processing. Using standardised protocols the average CV of serum profiles within- and between-run replicates has been evaluated and found to be around 10 % based on the variability of all detected peaks (more than 100 peaks per profile). The results show that this easy and fast method of using ZipTips, tested over a whole year, is more reproducible and more suitable for serum profile screening than magnetic bead-based methodologies which show higher variability and higher rates of rejected, low-quality spectra upon applying strict data quality control filters.
Volodya Vovk合作论文数Department of Computer Science5
Alex Gammerman合作论文数CLRC5
Knut Reinert合作论文数Freie Universit?0?1t Berlin;Institut f??r Informatik1