BACKGROUND:The annual incidence of pulmonary nodules is estimated at 1.57 million. Guidelines recommend using an initial assessment of nodule probability of malignancy (pCA). A previous study found that despite this recommendation, physicians did not follow guidelines. METHODS:Physician assessments (N = 337) and two previously validated risk model assessments of pretest probability of cancer were evaluated for performance in 337 patients with pulmonary nodules based on final diagnosis and compared. Physician-assessed pCA was categorized into low, intermediate, and high risk, and the next test ordered was evaluated. RESULTS:The prevalence of malignancy was 47% (n = 158) at 1 year. Physician-assessed pCA performed better than nodule prediction calculators (area under the curve, 0.85 vs 0.75; P < .001 and .78; P = .0001). Physicians did not follow indicated guidelines when selecting the next test in 61% of cases (n = 205). Despite recommendations for serial CT imaging in those with low pCA, 52% (n = 13) were managed more aggressively with PET imaging or biopsy; 12% (n = 3) underwent biopsy procedures for benign disease. Alternatively, in the high-risk category, the majority (n = 103 [75%]) were managed more conservatively. Stratified by diagnosis, 92% (n = 22) with benign disease underwent more conservative management with CT imaging (20%), PET scanning (15%), or biopsy (8%), although three had surgery (8%). CONCLUSIONS:Physician assessment as a means for predicting malignancy in pulmonary nodules is more accurate than previously validated nodule prediction calculators. Despite the accuracy of clinical intuition, physicians did not follow guideline-based recommendations when selecting the next diagnostic test. To provide optimal patient care, focus in the areas of guideline refinement, implementation, and dissemination is needed.
It is estimated that over 1.5 million lung nodules are detected annually in the United States. Most of these are benign but frequently undergo invasive and costly procedures to rule out malignancy. A risk predictor that can accurately differentiate benign and malignant lung nodules could be used to more efficiently route benign lung nodules to non-invasive observation by CT surveillance and route malignant lung nodules to invasive procedures. The majority of risk predictors developed to date are based exclusively on clinical risk factors, imaging technology or molecular markers. Assessed here are the relative performances of previously reported clinical risk factors and proteomic molecular markers for assessing cancer risk in lung nodules. From this analysis an integrated model incorporating clinical risk factors and proteomic molecular markers is developed and its performance assessed on a subset of 222 lung nodules, between 8mm and 20mm in diameter, collected in a previously reported prospective study. In this analysis it is found that the molecular marker is most predictive. However, the integration of clinical and molecular markers is superior to both clinical and molecular markers separately.CLINICAL TRIAL REGISTRATION:Registered at ClinicalTrials.gov (NCT01752101).
Evaluation of indeterminate pulmonary nodules is a complex challenge. Most are benign but frequently undergo invasive and costly procedures to rule out malignancy. A plasma protein classifier was developed that identifies likely benign nodules that can be triaged to CT surveillance to avoid unnecessary invasive procedures. The clinical utility of this classifier was assessed in a prospective-retrospective analysis of a study enrolling 475 patients with nodules 8-30 mm in diameter who had an invasive procedure to confirm diagnosis at 12 sites. Using this classifier, 32.0 % (CI 19.5-46.7) of surgeries and 31.8 % (CI 20.9-44.4) of invasive procedures (biopsy and/or surgery) on benign nodules could have been avoided. Patients with malignancy triaged to CT surveillance by the classifier would have been 24.0 % (CI 19.2-29.4). This rate is similar to that described in clinical practices (24.5 % CI 16.2-34.4). This study demonstrates the clinical utility of a non-invasive blood test for pulmonary nodules.
BACKGROUND:Current quantification methods for mass spectrometry (MS)-based proteomics either do not provide sufficient control of variability or are difficult to implement for routine clinical testing.RESULTS:We present here an integrated quantification (InteQuan) method that better controls pre-analytical and analytical variability than the popular quantification method using stable isotope-labeled standard peptides (SISQuan). We quantified 16 lung cancer biomarker candidates in human plasma samples in three assessment studies, using immunoaffinity depletion coupled with multiple reaction monitoring (MRM) MS. InteQuan outperformed SISQuan in precision in all three studies and tolerated a two-fold difference in sample loading. The three studies lasted over six months and encountered major changes in experimental settings. Nevertheless, plasma proteins in low ng/ml to low μg/ml concentrations were measured with a median technical coefficient of variation (CV) of 11.9% using InteQuan. The corresponding median CV using SISQuan was 15.3% after linear fitting. Furthermore, InteQuan surpassed SISQuan in measuring biological difference among clinical samples and in distinguishing benign versus cancer plasma samples.CONCLUSIONS:We demonstrated that InteQuan is a simple yet robust quantification method for MS-based quantitative proteomics, especially for applications in biomarker research and in routine clinical testing.
INTRODUCTION:Indeterminate pulmonary nodules (IPNs) lack clinical or radiographic features of benign etiologies and often undergo invasive procedures unnecessarily, suggesting potential roles for diagnostic adjuncts using molecular biomarkers. The primary objective was to validate a multivariate classifier that identifies likely benign lung nodules by assaying plasma protein expression levels, yielding a range of probability estimates based on high negative predictive values (NPVs) for patients with 8 to 30 mm IPNs. METHODS:A retrospective, multicenter, case-control study was performed using multiple reaction monitoring mass spectrometry, a classifier comprising five diagnostic and six normalization proteins, and blinded analysis of an independent validation set of plasma samples. RESULTS:The classifier achieved validation on 141 lung nodule-associated plasma samples based on predefined statistical goals to optimize sensitivity. Using a population based nonsmall-cell lung cancer prevalence estimate of 23% for 8 to 30 mm IPNs, the classifier identified likely benign lung nodules with 90% negative predictive value and 26% positive predictive value, as shown in our prior work, at 92% sensitivity and 20% specificity, with the lower bound of the classifier's performance at 70% sensitivity and 48% specificity. Classifier scores for the overall cohort were statistically independent of patient age, tobacco use, nodule size, and chronic obstructive pulmonary disease diagnosis. The classifier also demonstrated incremental diagnostic performance in combination with a four-parameter clinical model. CONCLUSIONS:This proteomic classifier provides a range of probability estimates for the likelihood of a benign etiology that may serve as a noninvasive, diagnostic adjunct for clinical assessments of patients with IPNs.
Each year, millions of pulmonary nodules are discovered by computed tomography and subsequently biopsied. Because most of these nodules are benign, many patients undergo unnecessary and costly invasive procedures. We present a 13-protein blood-based classifier that differentiates malignant and benign nodules with high confidence, thereby providing a diagnostic tool to avoid invasive biopsy on benign nodules. Using a systems biology strategy, we identified 371 protein candidates and developed a multiple reaction monitoring (MRM) assay for each. The MRM assays were applied in a three-site discovery study (n = 143) on plasma samples from patients with benign and stage IA lung cancer matched for nodule size, age, gender, and clinical site, producing a 13-protein classifier. The classifier was validated on an independent set of plasma samples (n = 104), exhibiting a negative predictive value (NPV) of 90%. Validation performance on samples from a nondiscovery clinical site showed an NPV of 94%, indicating the general effectiveness of the classifier. A pathway analysis demonstrated that the classifier proteins are likely modulated by a few transcription regulators (NF2L2, AHR, MYC, and FOS) that are associated with lung cancer, lung inflammation, and oxidative stress networks. The classifier score was independent of patient nodule size, smoking history, and age, which are risk factors used for clinical management of pulmonary nodules. Thus, this molecular test provides a potential complementary tool to help physicians in lung cancer diagnosis.
Colorectal cancer (CRC) is often curable and preventable using current screening modalities. Unfortunately, screening compliance remains low, partly due to patient dissatisfaction with faecal/endoscopic testing. Recent guidelines advise CRC screening should begin with risk stratification. A blood‐based test providing clinically actionable CRC risk information would likely improve screening compliance and enhance clinical decision making. We analyzed 196 gene expression profiles to select candidate CRC biomarkers. qRT‐PCR was performed on 642 samples to develop a 7‐gene biomarker panel using 112 CRC/120 controls (training set) and 202 CRC/208 controls (independent, blind test set). Panel performance characteristics and disease prevalence (0.7%) were then used to develop a scale assessing an individual's current risk of having CRC based on his/her gene signature. A 7‐gene panel (ANXA3, CLEC4D, LMNB1, PRRG4, TNFAIP6, VNN1 and IL2RB) discriminated CRC in the training set (area under the receiver‐operating‐characteristic curve (ROC AUC), 0.80; accuracy, 73%; sensitivity, 82%; specificity 64%). The independent blind test set confirmed performance (ROC AUC, 0.80; accuracy, 71%; sensitivity, 72%; specificity, 70%). Individual gene profiles were compared against the population results and used to calculate the current relative risk for CRC. We have developed a 7‐gene, blood‐based biomarker panel that can stratify subjects according to their current relative risk across a broad range in an average‐risk population. Across the continuous spectrum of risk as defined by the current relative risk scale, it is possible to identify clinically meaningful reference points that can assist patients and physicians in CRC screening decision making.
The analysis of tandem mass (MS/MS) data to identify and quantify proteins is hampered by the heterogeneity of file formats at the raw spectral data, peptide identification, and protein identification levels. Different mass spectrometers output their raw spectral data in a variety of proprietary formats, and alternative methods that assign peptides to MS/MS spectra and infer protein identifications from those peptide assignments each write their results in different formats. Here we describe an MS/MS analysis platform, the Trans-Proteomic Pipeline, which makes use of open XML file formats for storage of data at the raw spectral data, peptide, and protein levels. This platform enables uniform analysis and exchange of MS/MS data generated from a variety of different instruments, and assigned peptides using a variety of different database search programs. We demonstrate this by applying the pipeline to data sets generated by ThermoFinnigan LCQ, ABI 4700 MALDI-TOF/TOF, and Waters Q-TOF instruments, and searched in turn using SEQUEST, Mascot, and COMET.
It is expected that the composition of the serum proteome can provide valuable information about the state of the human body in health and disease and that this information can be extracted via quantitative proteomic measurements. Suitable proteomic techniques need to be sensitive, reproducible, and robust to detect potential biomarkers below the level of highly expressed proteins, generate data sets that are comparable between experiments and laboratories, and have high throughput to support statistical studies. Here we report a method for high throughput quantitative analysis of serum proteins. It consists of the selective isolation of peptides that are N-linked glycosylated in the intact protein, the analysis of these now deglycosylated peptides by liquid chromatography electrospray ionization mass spectrometry, and the comparative analysis of the resulting patterns. By focusing selectively on a few formerly N-linked glycopeptides per serum protein, the complexity of the analyte sample is significantly reduced and the sensitivity and throughput of serum proteome analysis are increased compared with the analysis of total tryptic peptides from unfractionated samples. We provide data that document the performance of the method and show that sera from untreated normal mice and genetically identical mice with carcinogen-induced skin cancer can be unambiguously discriminated using unsupervised clustering of the resulting peptide patterns. We further identify, by tandem mass spectrometry, some of the peptides that were consistently elevated in cancer mice compared with their control littermates.
Quantitative protein profiling using the isotope-coded affinity tag (ICAT) method and tandem mass spectrometry (MS) enables the pair-wise comparison of protein expression levels in biological samples. A new version of the ICAT reagent with an acid-cleavable bond, which allows removal of the biotin moiety prior to MS and which utilizes (13)C substitution for (12)C in the heavy-ICAT reagent rather than (2)H (for (1)H) as in the original reagent, was investigated. We developed and validated an MS data acquisition strategy using this new reagent that results in an increased number of protein identifications per experiment, without losing the accuracy of protein quantification. This was achieved by following a single survey (precursor) ion scan and serial collision induced dissociations (CIDs) of four different precursor ions observed in the prior survey scan. This strategy is common to many high-performance liquid chromatography-electrospray ionization (HPLC-ESI)-MS shotgun proteomic strategies, but heretofore not to ICAT experiments. This advance is possible because the new ICAT reagent uses (13)C as the "heavy" element rather than (2)H, thus, eliminating the slight delay in retention time of ICAT-labeled "light" peptides on a C18-based HPLC separation that occurs with (2)H and (1)H. Analyses using this new scheme of an ICAT-labeled trypsin-digested six protein mixture as well as a tryptic digest of a total yeast lysate, indicated that about two times more proteins were identified in a single analysis, and that there was no loss in accuracy of quantification.
There is an increasing interest in the quantitative proteomic measurement of the protein contents of substantially similar biological samples, e. g. for the analysis of cellular response to perturbations over time or for the discovery of protein biomarkers from clinical samples. Technical limitations of current proteomic platforms such as limited reproducibility and low throughput make this a challenging task. A new LC-MS-based platform is able to generate complex peptide patterns from the analysis of proteolyzed protein samples at high throughput and represents a promising approach for quantitative proteomics. A crucial component of the LC-MS approach is the accurate evaluation of the abundance of detected peptides over many samples and the identification of peptide features that can stratify samples with respect to their genetic, physiological, or environmental origins. We present here a new software suite, SpecArray, that generates a peptide versus sample array from a set of LC-MS data. A peptide array stores the relative abundance of thousands of peptide features in many samples and is in a format identical to that of a gene expression microarray. A peptide array can be subjected to an unsupervised clustering analysis to stratify samples or to a discriminant analysis to identify discriminatory peptide features. We applied the SpecArray to analyze two sets of LC-MS data: one was from four repeat LC-MS analyses of the same glycopeptide sample, and another was from LC-MS analysis of serum samples of five male and five female mice. We demonstrate through these two study cases that the SpecArray software suite can serve as an effective software platform in the LC-MS approach for quantitative proteomics.
In MS/MS experiments with automated precursor ion, selection only a fraction of sequencing attempts lead to the successful identification of a peptide. A number of reasons may contribute to this situation. They include poor fragmentation of the selected precursor ion, the presence of modified residues in the peptide, mismatches with sequence databases, and frequently, the concurrent fragmentation of multiple precursors in the same CID attempt. Current database search engines are incapable of correctly assigning the sequences of multiple precursors to such spectra. We have developed a search engine, ProbIDtree, which can identify multiple peptides from a CID spectrum generated by the concurrent fragmentation of multiple precursor ions. This is achieved by iterative database searching in which the submitted spectra are generated by subtracting the fragment ions assigned to a tentatively matched peptide from the acquired spectrum and in which each match is assigned a tentative probability score. Tentatively matched peptides are organized in a tree structure from which their adjusted probability scores are calculated and used to determine the correct identifications. The results using MALDI-TOF-TOF MS/MS data demonstrate that multiple peptides can be effectively identified simultaneously with high confidence using ProbIDtree.