Breast cancer manifests as multiple subtypes with distinct patient outcomes and treatment strategies. Here, we optimized proteomic analysis of Formalin-Fixed Paraffin-Embedded (FFPE) specimens from patients diagnosed with five breast cancer subtypes, luminal A, luminal B, Her2, triple negative (TNBC) and metaplastic breast cancers (MBC), and from disease-free individuals undergoing reduction mammoplasty (RM). We identified and quantified ∼6,000 protein groups (with >2 peptides per protein) with significant changes in over 26% of proteins comparing each cancer subtype with control RM. Stringent statistical filters allowed us to deeply mine 576 significant conserved protein changes shared by all subtypes and protein changes unique to each subtype. The most aggressive subtype, MBC, revealed exacerbated stromal stress responses, as illustrated by a collagenolytic extracellular matrix (ECM) and immune participation biased towards neutrophils and eosinophils. Immunostaining of breast tissue sections confirmed differences across subtypes, in particular, a strong upregulation of SERPINH1, neutrophil-specific myeloperoxidase and eosinophil cationic protein in MBC. In summary, we present deep proteomic, digitalized protein abundance profiles, generated from FFPE breast cancer tissues, that revealed significant changes in ECM and cellular proteins. Statement of Significance of the Study:This study is significant as it discovered deep proteomic signatures for the highly aggressive and malignant metaplastic breast cancer (MCB) which is now considered a fifth subtype based upon its remarkable intra-tumoral heterogeneity that illustrates its unique cell plasticity. To efficiently analyze formalin-fixed paraffin-embedded (FFPE) breast tissues from patients with different breast cancer subtypes and disease-free individuals, we optimized a novel workflow in which we combined paraffinization and Folch extraction. We identified and confidently quantified ∼6,000 protein groups. We were able to find robust changes in extracellular matrix (ECM) with cancer, even though no ECM enrichments were performed. Interestingly, despite the relatively small human cohort size (42 patients), distinct protein signatures emerged throughout all cancer subtypes - common and unique - with remarkable statistical significance for many cancer-relevant proteins and pathways. This Pilot study indicates the hypothesis that an altered stroma can dictate epithelial tumor cell fate. We also observed that MBC was characterized by an especially immunosuppressed tumor environment. We do note the limitation of the relatively small cohort size of our study, and in the future additional patient cohorts will be needed to further validate our findings.
The global scientific response to COVID 19 highlighted the urgent need for increased throughput and capacity in bioanalytical laboratories, especially for the precise quantification of proteins that pertain to health and disease. Acoustic ejection mass spectrometry (AEMS) represents a much-needed paradigm shift for ultra-fast biomarker screening. Here, a quantitative AEMS assays is presented, employing peptide immunocapture to enrich (i) 10 acute phase response (APR) protein markers from plasma, and (ii) SARS-CoV-2 NCAP peptides from nasopharyngeal swabs. The APR proteins were quantified in 267 plasma samples, in triplicate in 4.8 h, with %CV from 4.2% to 10.5%. SARS-CoV-2 peptides were quantified in triplicate from 145 viral swabs in 10 min. This assay represents a 15-fold speed improvement over LC-MS, with instrument stability demonstrated across 10,000 peptide measurements. The combination of speed from AEMS and selectivity from peptide immunocapture enables ultra-high throughput, reproducible quantitative biomarker screening in very large cohorts. There is an increasing demand for high sample throughput beyond traditional analytical techniques. Here the authors show that Ultra-high throughput assays, based on the integration of Acoustic Ejection Mass Spectrometry (AEMS) and peptide immunocapture provide precise quantification of inflammation and SARS-CoV-2 proteins with 15x sample throughput compared to LC-MS.
The human immunodeficiency virus (HIV) integrates into the host genome forming latent cellular reservoirs that are an obstacle for cure or remission strategies. Viral transcription is the first step in the control of latency and depends upon the hijacking of the host cell RNA polymerase II (Pol II) machinery by the 5’ HIV LTR. Consequently, “block and lock” or “shock and kill” strategies for an HIV cure depend upon a full understanding of HIV transcriptional control. The HIV trans-activating protein, Tat, controls HIV latency as part of a positive feed-forward loop that strongly activates HIV transcription. The recognition of the TATA box and adjacent sequences of HIV essential for Tat trans-activation (TASHET) of the core promoter by host cell pre-initiation complexes of HIV (PICH) has been shown to be necessary for Tat trans-activation, yet the protein composition of PICH has remained obscure. Here, DNA-affinity chromatography was employed to identify the mitotic deacetylase complex (MiDAC) as selectively recognizing TASHET. Using biophysical techniques, we show that the MiDAC subunit DNTTIP1 binds directly to TASHET, in part via its CTGC DNA motifs. Using co-immunoprecipitation assays, we show that DNTTIP1 interacts with MiDAC subunits MIDEAS and HDAC1/2. The Tat-interacting protein, NAT10, is also present in HIV-bound MiDAC. Gene silencing revealed a functional role for DNTTIP1, MIDEAS, and NAT10 in HIV expression in cellulo. Furthermore, point mutations in TASHET that prevent DNTTIP1 binding block the reactivation of HIV by latency reversing agents (LRA) that act via the P-TEFb/7SK axis. Our data reveal a key role for MiDAC subunits DNTTIP1, MIDEAS, as well as NAT10, in Tat-activated HIV transcription and latency. DNTTIP1, MIDEAS and NAT10 emerge as cell cycle-regulated host cell transcription factors that can control activated HIV gene expression, and as new drug targets for HIV cure strategies.
Untargeted metabolomics based on reverse phase LC-MS (RPLC-MS) plays a crucial role in biomarker discovery across physiological and disease states. Standardizing the development process of untargeted methods requires paying attention to critical factors that are under discussed or easily overlooked, such as injection parameters, performance assessment, and matrix effect evaluation. In this study, we developed an untargeted metabolomics method for plasma and fecal samples with the optimization and evaluation of these factors. Our results showed that optimizing the reconstitution solvent and sample injection amount was critical for achieving the balance between metabolites coverage and signal linearity. Method validation with representative stable isotopically labeled standards (SILs) provided insights into the analytical performance evaluation of our method. To tackle the issue of the matrix effect, we implemented a postcolumn infusion (PCI) approach to monitor the overall absolute matrix effect (AME) and relative matrix effect (RME). The monitoring revealed distinct AME and RME profiles in plasma and feces. Comparing RME data obtained for SILs through postextraction spiking with those monitored using PCI compounds demonstrated the comparability of these two methods for RME assessment. Therefore, we applied the PCI approach to predict the RME of 305 target compounds covered in our in-house library and found that targets detected in the negative polarity were more vulnerable to the RME, regardless of the sample matrix. Given the value of this PCI approach in identifying the strengths and weaknesses of our method in terms of the matrix effect, we recommend implementing a PCI approach during method development and applying it routinely in untargeted metabolomics.
Acute kidney injury (AKI) manifests as a major health concern, particularly for the elderly. Understanding AKI-related proteome changes is critical for prevention and development of novel therapeutics to recover kidney function and to mitigate the susceptibility for recurrent AKI or development of chronic kidney disease. In this study, mouse kidneys were subjected to ischemia-reperfusion injury, and the contralateral kidneys remained uninjured to enable comparison and assess injury-induced changes in the kidney proteome. A ZenoTOF 7600 mass spectrometer was optimized for data-independent acquisition (DIA) to achieve comprehensive protein identification and quantification. Short microflow gradients and the generation of a deep kidney-specific spectral library allowed for high-throughput, comprehensive protein quantification. Upon AKI, the kidney proteome was completely remodeled, and over half of the 3945 quantified protein groups changed significantly. Downregulated proteins in the injured kidney were involved in energy production, including numerous peroxisomal matrix proteins that function in fatty acid oxidation, such as ACOX1, CAT, EHHADH, ACOT4, ACOT8, and Scp2. Injured kidneys exhibited severely damaged tissues and injury markers. The comprehensive and sensitive kidney-specific DIA-MS assays feature high-throughput analytical capabilities to achieve deep coverage of the kidney proteome, and will serve as useful tools for developing novel therapeutics to remediate kidney function.
The human immunodeficiency virus (HIV) integrates into the host genome forming latent cellular reservoirs that are an obstacle for cure or remission strategies. Viral transcription is the first step in the control of latency and depends upon the hijacking of the host cell RNA polymerase II (Pol II) machinery by the 5’ HIV LTR. Consequently, “block and lock” or “shock and kill” strategies for an HIV cure depend upon a full understanding of HIV transcriptional control. The HIV trans-activating protein, Tat, controls HIV latency as part of a positive feed-forward loop that strongly activates HIV transcription. The recognition of the T ATA box and a djacent s equences of H IV e ssential for T at trans -activation (TASHET) of the core promoter by host cell p re-initiation c omplexes of H IV (PICH) has been shown to be necessary for Tat trans -activation, yet the protein composition of PICH has remained obscure. Here DNA-affinity chromatography was employed to identify the mitotic deacetylase complex (MiDAC) as selectively recognizing TASHET. Using biophysical techniques, we show that the MiDAC subunit DNTTIP1 binds directly to TASHET, in part via its CTGC DNA motifs. Using co-immunoprecipitation assays, we show that DNTTIP1 interacts with MiDAC subunits MIDEAS and HDAC1/2. The Tat-interacting protein, NAT10, is also present in HIV-bound MiDAC. Gene silencing revealed a functional role for DNTTIP1, MIDEAS, and NAT10 in HIV expression in cellulo . Furthermore, point mutations in TASHET that prevent DNTTIP1 binding block the reactivation of HIV by latency reversing agents (LRA) that act via the P-TEFb/7SK axis in a model of latency. Our data reveal a key role for MiDAC subunits DNTTIP1, MIDEAS, as well as NAT10, in Tat-activated HIV transcription and latency. DNTTIP1, MIDEAS and NAT10 emerge as cell cycle-regulated host cell transcription factors that can control HIV latency, and as new drug targets for HIV cure strategies. Author summary Latent HIV integrated within the host cell genome poses a major problem for viral eradication. The reactivation of latent HIV depends on host cell transcription factors that are hijacked by the 5’ long terminal repeat (LTR) region of the HIV genome to produce viral RNA. At the heart of the LTR of HIV lies a particularly crucial DNA region named the core promoter that is specifically required for the reactivation of HIV by a viral protein named Tat. A significant body of work over more than 30 years has established the specific requirement for the HIV core promoter in Tat’s control of HIV latency, but the underlying molecular mechanisms have remained elusive. Here, we identify host cell transcription factors that bind selectively to the HIV core promoter to control HIV gene expression. Our data reveal three human proteins that act in a complex to reactivate latent HIV, including one that directly recognizes the HIV core promoter, one that regulates chromatin, and a third that binds to the HIV Tat protein. Our data fill a significant and long-standing gap in the understanding of latency and identify new potential drug targets for HIV cure strategies.
The methodology of data-independent acquisition (DIA) within mass spectrometry (MS) was developed into a method of choice for quantitative proteomics, to capture the depth and dynamics of biological systems, and to perform large-scale protein quantification. DIA provides deep quantitative proteome coverage with high sensitivity, high quantitative accuracy, and excellent acquisition-to-acquisition reproducibility. DIA workflows benefited from the latest advancements in MS instrumentation, acquisition/isolation schemes, and computational algorithms, which have further improved data quality and sample throughput. This powerful DIA-MS scan type selects all precursor ions contained in pre-determined isolation windows, and systematically fragments all precursor ions from each window by tandem mass spectrometry, subsequently covering the entire precursor ion m/z range. Comprehensive proteolytic peptide identification and label-free quantification are achieved post-acquisition using spectral library-based or library-free approaches. To celebrate the > 10 years of success of this quantitative DIA workflow, we interviewed some of the scientific leaders who have provided crucial improvements to DIA, to the quantification accuracy and proteome depth achieved, and who have explored DIA applications across a wide range of biology. We discuss acquisition strategies that improve specificity using different isolation schemes, and that reduce complexity by combining DIA with sophisticated chromatography or ion mobility separation. Significant leaps forward were achieved by evolving data processing strategies, such as library-free processing, and machine learning to interrogate data more deeply. Finally, we highlight some of the diverse biological applications that use DIA-MS methods, including large-scale quantitative proteomics, post-translational modification studies, single-cell analysis, food science, forensics, and small molecule analysis.
ABSTRACT Protein post-translational modifications (PTMs) are crucial and dynamic players in a large variety of cellular processes and signaling, and proteomic technologies have emerged as the method of choice to profile PTMs. However, these analyses remain challenging due to potential low PTM stoichiometry, the presence of multiple PTMs per proteolytic peptide, PTM site localization of isobaric peptides, and labile PTM groups that lead to neutral losses. Collision-induced dissociation (CID) is commonly used for to characterize PTMs, but the application of collision energy can lead to neutral losses and incomplete peptide sequencing for labile PTM groups. In this study, we compared CID to an alternative fragmentation, electron activated dissociation (EAD), operated on a recently introduced fast-acquisition quadrupole-time-of-flight (QqTOF) mass spectrometer. We analyzed a series of synthetic modified peptides, featuring phosphorylated, succinylated, malonylated, and acetylated peptides. We performed targeted, quantitative parallel reaction monitoring (PRM or MRM HR ) assays to assess the performances of EAD to characterize, site-localize and quantify peptides with labile modifications. The tunable EAD kinetic energy allowed the preservation of labile modifications and provided better peptide sequence coverage with strong PTM-site localization fragment ions. Zeno trap activation provided significant MS/MS sensitivity gains by an average of 6–11-fold for EAD analyses, regardless of modification type. Evaluation of the quantitative EAD PRM workflows revealed high reproducibility with coefficients of variation of typically ∼2%, as well as very good linearity and quantification accuracy. This novel workflow, combining EAD and Zeno trap, offers confident, accurate, and robust characterization and quantification of PTMs.
Modern biomarker and translational research as well as personalized health care studies rely heavily on powerful omics' technologies, including metabolomics and lipidomics. However, to translate metabolomics and lipidomics discoveries into a high-throughput clinical setting, standardization is of utmost importance. Here, we compared and benchmarked a quantitative lipidomics platform. The employed Lipidyzer platform is based on lipid class separation by means of differential mobility spectrometry with subsequent multiple reaction monitoring. Quantitation is achieved by the use of 54 deuterated internal standards and an automated informatics approach. We investigated the platform performance across nine laboratories using NIST SRM 1950–Metabolites in Frozen Human Plasma, and three NIST Candidate Reference Materials 8231–Frozen Human Plasma Suite for Metabolomics (high triglyceride, diabetic, and African-American plasma). In addition, we comparatively analyzed 59 plasma samples from individuals with familial hypercholesterolemia from a clinical cohort study. We provide evidence that the more practical methyl-tert-butyl ether extraction outperforms the classic Bligh and Dyer approach and compare our results with two previously published ring trials. In summary, we present standardized lipidomics protocols, allowing for the highly reproducible analysis of several hundred human plasma lipids, and present detailed molecular information for potentially disease relevant and ethnicity-related materials.
We reported and evaluated a microflow, single-shot, short gradient SWATH MS method intended to accelerate the discovery and verification of protein biomarkers in preclassified clinical specimens. The method uses a 15 min gradient microflow-LC peptide separation, an optimized SWATH MS window configuration, and OpenSWATH software for data analysis. We applied the method to a cohort containing 204 FFPE tissue samples from 58 prostate cancer patients and 10 benign prostatic hyperplasia patients. Altogether we identified 27,975 proteotypic peptides and 4037 SwissProt proteins from these 204 samples. Compared to a reference SWATH method with a 2 h gradient, we found 3800 proteins were quantified by the two methods on two different instruments with relatively high consistency (r = 0.77). The accelerated method consumed only 17% instrument time, while quantifying 80% of proteins compared to the 2 h gradient SWATH. Although the missing value rate increased by 20%, batch effects reduced by 21%. 75 deregulated proteins measured by the accelerated method were selected for further validation. A shortlist of 134 selected peptide precursors from the 75 proteins were analyzed using MRM-HR, and the results exhibited high quantitative consistency with the 15 min SWATH method (r = 0.89) in the same sample set. We further verified the applicability of these 75 proteins in separating benign and malignant tissues (AUC = 0.99) in an independent prostate cancer cohort (n = 154). Altogether, the results showed that the 15 min gradient microflow SWATH accelerated large-scale data acquisition by 6 times, reduced batch effect by 21%, introduced 20% more missing values, and exhibited comparable ability to separate disease groups.
The combination of microflow LC with SWATH Acquisition for large scale quantitative proteomics studies is becoming increasingly more widespread, due to the improved robustness and throughput obtained relative to the traditional nanoflow LC approach. High quality quantitative datasets have been generated using a standard 1 hour gradient, demonstrating large numbers of proteins quantified routinely.The impact of gradient length on protein identification results using data dependent acquisition (DDA) was also explored previously and summarized here. Next, an exploration into the impact of gradient length on proteins quantified using data independent acquisition (DIA) was undertaken to provide researchers expanded workflow options with microflow SWATH acquisition. Using microflow liquid chromatography gradients as short as five minutes (total run time <15 mins), SWATH acquisition parameters were first optimized then used on multiple TripleTOF® 6600 systems to study impact of shortened separations on number of proteins quantified.Testing on multiple instruments with different complex matrices, it was observed that the number of proteins quantified using SWATH acquisition for the fastest gradients was over 1300 proteins for multiple complex matrices from 1 μg protein load. When combined with larger ion libraries, over 2100 proteins were identified and quantified in a 5 min gradient from a complex HEK digest (1 μg). The combination of microflow LC with the high MS/MS acquisition rates of the TripleTOF® 6600 system enables high‐throughput yet comprehensive analysis for proteomics samples, at rates approaching 100 samples per day.
Abstract BACKGROUND Discovery and verification of protein biomarkers in clinical specimens using mass spectrometry are inherently challenging and resource-intensive. METHODS Formalin-fixed paraffin-embedded tissue-biopsy samples from a prostate cancer patient cohort (PCZA, n = 68) were processed in triplicate using pressure cycling technology, followed by microflow LC SWATH® analysis with different gradients and window schemes. Potential protein biomarker candidates were prioritized using random forest analysis and evaluated by receiver operating characteristic curve analysis. Selected proteins were verified with a targeted MRMHR assay using the 15 min microflow LC strategy on a second prostate cancer cohort (PCZB, n = 54). Potential biomarkers were further verified using TMA on a third cohort (PCZD, n = 100). RESULTS We developed and optimized a 15-min microflow LC approach coupled with microflow SWATH MS. Application of the optimal 15 min and conventional 120 min LC gradient scheme using samples from the PCZA cohort led to quantification of 3,800 proteins in both methods with high quantitative correlation (r = 0.77). MRMHR verification of 154 prioritized proteins showed high quantitative consistency with the 15 min SWATH data (r = 0.89). Separation of benign and malignant tissues achieved precision (AUC = 0.99). ECHS1 was further verified in a third cohort PCZD successfully, agreeing with RNAseq data from the TCGA in a different cohort (n=549). Our methods enables practical proteomic analysis of 204 tissue samples within 5 working days. CONCLUSION Single-shot, short gradient SWATH-MS coupled with MRMHR is both practical and effective in discovering and verifying protein biomarkers in clinical specimens.
The complexity of a proteomics sample after digestion is extremely high requiring that extensive fractionation is done to deeply interrogate the proteome. The key goal is to spread the peptides out across fractions such that when each is analyzed by LC-MS/MS, the mass spectrometer has time to collect high quality MS/MS spectra on as many peptides as possible. Typically when more fractions are collected, more protein identifications are obtained. The downside is that the more fractions collected to increase depth of coverage, the more instrument time is needed to analyze all the fractions.
Protein citrullination (or deimination), an irreversible post-translational modification, has been implicated in several physiological and pathological processes, including gene expression regulation, apoptosis, rheumatoid arthritis, and Alzheimer's disease. Several research studies have been carried out on citrullination under many conditions. However, until now, challenges in sample preparation and data analysis have made it difficult to confidently identify a citrullinated protein and assign the citrullinated site. To overcome these limitations, we generated a mouse hyper-citrullinated spectral library and set up coordinates to confidently identify and validate citrullinated sites. Using this workflow, we detect a four-fold increase in citrullinated proteome coverage across six mouse organs compared with the current state-of-the art techniques. Our data reveal that the subcellular distribution of citrullinated proteins is tissue-type-dependent and that citrullinated targets are involved in fundamental physiological processes, including the metabolic process. These data represent the first report of a hyper-citrullinated library for the mouse and serve as a central resource for exploring the role of citrullination in this organism.
Mass spectrometry has transformed quantitative analysis to become the method of choice for many assays. More recently, LC-MS/MS has revolutionized quantitative bioanalysis. While single MS filtering offers advantages over non-mass selective techniques, the use of tandem mass spectrometry (MS/MS, or MS2) eliminates interferences and results in a dramatic increase in selectivity which yields a very low baseline, excellent limits of quantification, and very good linearity. As a result, the Multiple Reaction Monitoring (MRM) experiment performed on triple quadrupole mass spectrometers has become the technique of choice for highly sensitive and selective quantification in biological matrices.
In recent years, the biochemical study of lipids has transformed from slow multi-dimensional chromatographic separations and chemical derivatization strategies to higher throughput analysis by mass spectrometry. Advances in mass spectrometry have enabled in-depth lipidomic analyses with unparalleled qualitative and quantitative sensitivity. However, unambiguous identification and quantitation of lipid molecular species in total lipid extracts has proven to be difficult, primarily due to isobaric overlapping isobaric and isomeric species. There are greater than 100,000 lipid molecular species present in a typical biological lipid extract that occupy a narrow mass range (~400-1100 amu), making such overlap a significant problem.
Immunoglobulin G molecules have become attractive as targeted therapeutic proteins, due to their high specificity and long circulation time. Glycosylation patterns determine the stability and bio-disposition of these recombinant protein drugs in vivo, as well as the efficacy, folding, binding affinity, specificity and pharmacokinetic properties. Therefore, a complete characterization of the biotherapeutic IgG glycosylation is desirable.