The accessory gene regulator (agr) quorum sensing system in Staphylococcus aureus plays a prominent role in acute infections caused by this common pathogen. In contrast, ATP synthase affects the ability of S. aureus to establish chronic infections and persist as small colony variants. In this study, we report that a synthetic small molecule, CP-20, is capable of dually inhibiting agr and ATP synthase. We demonstrate CP-20’s ability to counteract key virulence phenotypes associated with both pathways, including hemolysis, biofilm formation, and antibiotic resistance. We also show that CP-20 facilitates more efficient clearance of S. aureus by murine macrophages than genetic inactivation of either agr or ATP synthase alone, suggesting that dual inhibition may have an additive effect on ablating intracellular survival. This work demonstrates the utility of a single chemical modulator to attenuate S. aureus pathogenesis via two non-essential pathways and suggests a pathway toward non-bactericidal virulence control.
Peptide mapping is a critical tool for characterizing biotherapeutic proteins and is essential for the development of monoclonal antibody drugs. Here we describe a new direct infusion technology that streamlines peptide mapping data collection and analysis, accelerating the method by up to 100-fold. This method, which we term RaPiD-mAb-MS, combines high-throughput plate-based sample preparation with direct infusion mass spectrometry analysis. RaPiD-mAb-MS allows analysis of 96 samples within ∼ 1.5 to 2 hours, routinely achieves >95% sequence coverage, and has been successfully applied to 28 unique antibodies and over 2,000 samples. Here we demonstrate that RaPiD-mAb-MS detects and quantifies oxidation, deamidation, isomerization, glycosylation, and sequence variants with results comparable to conventional LC-MS based methods in a fraction of the time. Further, by eliminating chromatography, data analysis is greatly streamlined and simplified. By allowing for the collection of ∼ 1,000 peptide maps per day, RaPiD-mAb-MS is positioned to accelerate all phases of antibody-based drug discovery & development and sets the stage for collection of massive datasets that would allow artificial intelligent prediction of optimal antibody variants and formulations.
Isotope distribution prediction is an important part of mass spectrometry data analysis. A variety of strategies have been developed, including brute-force polynomial methods and Fourier-transform (FT) convolution methods. Here, we present a novel neural network (NN) approach to isotope prediction. These NN-based tools are distributed in a new package, IsoGen, alongside FT-based tools. We show that NNs perform as well as existing approaches when predicting isotope distributions for molecules with natural isotope ratios. We then demonstrate the capability of NNs for transfer learning, a method by which existing models that have been trained to perform one task can be reused as a starting point to train models to perform a second task. Here, a NN that has been trained to predict isotope distributions for molecules with natural isotope ratios can be retrained to predict distributions for molecules with perturbed isotope ratios. We also demonstrate that these training distributions do not need to be produced theoretically but may be extracted from experimental data where the underlying isotopic composition is not known. Overall, NNs provide a robust method for isotope prediction that can be extended to applications where the isotope distributions can be measured but not easily predefined.
The dehydroamino acids (DHAAs) dehydroalanine and dehydrobutyrine are formed in proteins posttranslationally from serine and threonine, respectively. They contain an electrophilic alkene, which reacts readily with nucleophiles by Michael addition, giving rise to a variety of chemical modifications as well as protein crosslinks. Recent work has shown DHAAs and their modifications to be highly prevalent in protein aggregates isolated from brains of Alzheimer's Disease (AD) patients. As there is no human enzyme known to catalyze the formation of DHAAs, this raised the question of how they are produced. We show here that metal oxides, particularly the iron oxides goethite and hematite, efficiently catalyze the formation of DHAAs in model phosphopeptides in vitro under physiologically relevant conditions. Iron oxides are known to accumulate in the human brain with age, and this deposition is enhanced in neurodegenerative disorders like AD. The DHAA formation by heterogeneous catalysis reported here provides a previously unknown mechanism linking iron deposition to protein crosslinking in AD and potentially other neurodegenerative diseases where iron accumulation and protein aggregation co-occur.
Proximity labeling (PL) can identify endogenous proteins in specific subcellular regions. APEX2 (enhanced ascorbate peroxidase 2) enables PL with high temporal resolution by oxidizing biotin phenol (BP) into a radical that promiscuously tags nearby proteins, followed by the isolation and identification of the tagged biomolecules. However, the BP radical has a relatively large diffusion radius, limiting spatial resolution. Replacing the phenol in BP with nitrophenol (NP) could shorten the labeling radius by shortening the radical lifetime, but existing PL heme peroxidases cannot efficiently oxidize substrates with high redox potentials. We report the ultrahigh-throughput directed evolution of APOX (APEX with high redOX potential), a quadruple mutant of APEX2 with a higher reduction potential. Using APOX with a membrane-permeable alkyne-NP probe, we demonstrate PL in living mammalian cells, including proteomic mapping with excellent subcellular compartment specificity.
Alzheimer's disease (AD) is characterized by the accumulation of protein aggregates, which are thought to be influenced by posttranslational modifications (PTMs). Dehydroamino acids (DHAAs) are rarely observed PTMs that contain an electrophilic alkene capable of forming protein-protein crosslinks, which may lead to protein aggregation. We report here the discovery of DHAAs in the protein aggregates from AD, constituting an unknown and previously unsuspected source of extensive proteomic complexity. We used mass spectrometry-based proteomics to discover 404 sites of DHAA formation in 171 proteins from protein aggregate-enriched human brain samples, 6-fold more sites than observed in the soluble protein fractions. The DHAA modifications are observed both directly and in the form of conjugates after reacting with abundant cellular nucleophiles or crosslinking to nucleophilic amino acid residues. We report 11 such crosslinks, including three in the Tau protein, which are 10-fold more abundant in AD samples compared with age-matched controls. Many of the proteins found to contain DHAAs and their conjugates are involved in protein aggregation or pathways dysregulated in AD. DHAAs are prevalent modifications in the AD brain proteome and give rise to protein crosslinks that may contribute to protein aggregation.
Long noncoding RNAs (lncRNAs) exert regulatory functions in a wide spectrum of biological contexts, and certain regulatory functions involve the formation of RNA-protein complexes. Discovering the structure and function of these complexes may unveil important functional insights. The DDX41 gene encoding the DEAD-box RNA helicase 41 protein (DDX41) is subject to extensive germline genetic variation, and certain variants create a predisposition to develop myelodysplastic syndrome and acute myeloid leukemia. While the importance of DDX41 for the control of hematopoiesis is established, many questions remain regarding the mechanisms of how DDX41 functions in hematopoietic stem and progenitor cells. Previously, we identified a DDX41-regulated lncRNA, growth-arrest-specific 5 (Gas5). As the Gas5 function in hematopoiesis is unknown, we analyzed the protein interactors of Gas5 lncRNA using HyPR-MS (hybridization purification of RNA-protein complexes, followed by mass spectrometry). A total of 303 proteins were identified as Gas5 lncRNA interactors, five of which were experimentally validated as Gas5 lncRNA interactors by RNA immunoprecipitation qPCR (RIP-qPCR) analysis. The identification of protein interactors with a DDX41-regulated lncRNA establishes a foundation on which to guide future mechanistic and biological studies.
How RIF1 (RAP1 interacting factor 1) fulfills its diverse roles in DNA double-strand break repair, DNA replication, and nuclear organization remains elusive. Here, we show that alternative splicing of a cassette exon (Ex32) encoding a Ser/Lys-rich cassette in the RIF1 C-terminal domain (CTD) gives rise to RIF1-Long (RIF1-L) and RIF1-Short (RIF1-S) isoforms with different functional characteristics. We demonstrate that RIF1-Ex32 splice-in is mediated by an exonic splicing enhancer that is recognized by the serine and arginine rich splicing factor 1 (SRSF1) and antagonized by SRSF3 and SRSF7. Exposure to DNA damage inhibited Ex32 splice-in, potentiated the association of SRSF3 and SRSF7 with RIF1 pre-mRNA, and caused an increase in RIF1-S protein expression, which was also observed across a diverse set of primary cancers. Isoform-specific proteomic analyses revealed RIF1-L preferentially associated with mediator of DNA damage checkpoint 1 (MDC1) and sustained MDC1 focus formation to a greater extent than RIF1-S. We further show that the Ser/Lys-rich cassette stabilized a novel phase separation activity of the RIF1 CTD and enhanced RIF1-L chromatin retention, which was reversed by cyclin-dependent kinase 1-dependent phosphorylation of the RIF1 CTD in response to G2 DNA damage checkpoint inhibition. These combined findings suggest DNA damage-dependent RIF1 alternative splicing contributes to RIF1 functional diversification in genome protection.
In both native and engineered tissues, the extracellular matrix (ECM) supports and regulates nearly all aspects of cellular pathophysiology, and in response, cells extensively remodel their surrounding extracellular environments through new ECM protein deposition. Understanding this intricate bi-directional cell-ECM interaction is key to tissue engineering, but it remains challenging to investigate. This is partly due to the limited sensitivity of conventional proteomics to capture low-abundance newly synthesized ECM (newsECM). This study presents a glycosylation-enabled, chemoselective strategy to label, enrich, and characterize newsECM proteins with augmented specificity and sensitivity. Applying newsECM profiling to bioengineered tumor tissues, either built upon decellularized ECM materials (dECM-tumor) or as ECM-free tumoroids, revealed distinct ECM synthesis patterns. Tumor cells cultured within dECM scaffold present elevated ECM remodeling activities, mediated by augmented digestion of pre-existing ECM coupled with upregulated synthesis of tumor-associated ECM components. These findings highlight the sensitivity of newsECM profiling to capture remodeling events that are otherwise under-represented by bulk proteomics and underscore the significance of dECM support for enabling native-like tumor cell behaviors. The newsECM profiling described here is anticipated to be applicable to a wide range of engineered tissue models and pathophysiological processes to deliver fundamental insights regarding the mutual cell-ECM crosstalk.
The extracellular matrix (ECM), present in nearly all tissues, provides extensive support to resident cells through structural, biomechanical, and biochemical means, and in return the ECM undergoes constant remodeling from interacting cells to adapt to the evolving tissue states. Bioengineered 3D tissues, commonly known as cell-ECM composites, are robust model systems to recapitulate and investigate native pathophysiology. Key to this engineered morphogenesis process are the intricate cell-ECM interactions reflected by how cells respond to and thereby modulate their surrounding microenvironments through their ongoing ECM secretome. However, investigating ECM-regulated new ECM production has been challenging due to the proteomic background from the pre-existing biomaterial ECM. To address this hindrance, here we present a chemoselective strategy to label, enrich, and characterize newly synthesized ECM (newsECM) proteins produced by resident cells, allowing distinction from the pre-existing ECM background. Applying our analytical pipeline to bioengineered tumor tissues, either built upon decellularized ECM (dECM-tumors) or as ECM-free tumor spheroids (tumoroids), we observed distinct ECM synthesis patterns that were linked to their extracellular environments. Tumor cells responded to the dECM presence with elevated ECM remodeling activities, mediated by augmented digestion of pre-existing ECM coupled with upregulated synthesis of tumor-associated ECM. Our findings highlight the sensitivity of newsECM profiling to capture remodeling events that are otherwise under-represented by bulk proteomics and underscore the significance of dECM support for enabling native-like tumor cell behaviors. We anticipate the described newsECM analytical pipeline to be broadly applicable to other tissue-engineered systems to probe ECM-regulated ECM synthesis and remodeling, both fundamental aspects of cell-ECM crosstalk in engineered tissue morphogenesis.
Histone post-translational modifications (PTMs) are crucial to eukaryotic genome regulation, with a range of reported functions and mechanisms of action. Though often studied individually, it has long been recognized that the modifications function by combinatorial synergy or antagonism. Interplay may involve PTMs on the same histone, within the same nucleosome (containing a histone octamer), or between nucleosomes in higher-order chromatin. Given this, the field must distinguish ever greater complexity, and the context in which it is studied, with brevity and precision. The proteoform was introduced to define individual forms of a protein by sequence and PTMs, followed by the nucleoform to describe the particular gathering of histones within an individual nucleosome. There is now a need to define specific forms of these entities in prose while providing space for experimental nuance. To this end, we introduce a nomenclature that can express discrete PTMs, proteoforms, nucleoforms, or situations where defined PTMs exist in an uncertain context. Though specifically designed for the chromatin field, adaptions of the framework could be used to describe-and thus dissect-how proteoforms are configured in functionally distinct complexes across biology.
Electrospray ionization (ESI) mass spectrometry is an essential technique for chemical analysis in a range of fields. In ESI, analytes can produce multiple charge states, which must be correctly assigned for identification. Existing approaches to charge state assignment can suffer from limited accuracy or poor speed. Here, we developed a fast neural network to perform isotopic envelope charge assignment. The performance of our algorithm, IsoDec, was demonstrated on top-down proteomics spectra collected on diverse instruments. On these highly complex individual spectra, we found that IsoDec correctly assigns more features compared to existing software tools while simultaneously providing improved speed and accuracy. Importantly, this performance enhancement stems directly from the neural network charge assignment approach and not simply from improved scoring and filtering of isotopic envelopes. Finally, when applied to large top-down proteomics data sets, we discovered that database searching of the IsoDec deconvolution output produces proteoform-spectrum matches with a better combination of coverage and accuracy. Overall, IsoDec provides a compelling demonstration of the potential of lightweight neural networks in mass spectrometry data analysis for diverse applications.
The goal of proteomics is to identify and quantify peptides and proteins within a biological sample. Almost all algorithms for the identification of peptides in LC-MS/MS data employ two steps: peptide/spectrum matching and peptide-identity-propagation (PIP), also known as match-between-runs. PIP can routinely account for up to 40% of all results, with that proportion rising as high as 75% in single-cell proteomics. Unlike peptide identities derived through peptide/spectrum matches, for which error estimation has been strictly enforced for decades, peptide identities derived through PIP have not historically been subject to statistical evaluation. As an indispensable component of label-free quantification, PIP needs a statistically rigorous method for estimating its false-discovery rate (FDR). We present a method for FDR control of PIP, called PIP-ECHO, and devise a rigorous protocol for evaluating FDR control of any PIP method. Using three different benchmark data sets, we evaluate PIP-ECHO alongside the PIP procedures implemented by FlashLFQ, IonQuant, and MaxQuant. These analyses show that only PIP-ECHO can accurately control the FDR of PIP at 1% across all data sets. When analyzing a spike-in data set, PIP-ECHO increases both the accuracy and sensitivity of differential expression analysis, yielding substantially more differentially abundant proteins than either MaxQuant or IonQuant.
Motivation:Studying protein isoforms is an essential step in biomedical research; at present, the main approach for analyzing proteins is via bottom-up mass spectrometry proteomics, which return peptide identifications, that are indirectly used to infer the presence of protein isoforms. However, the detection and quantification processes are noisy; in particular, peptides may be erroneously detected, and most peptides, known as shared peptides, are associated to multiple protein isoforms. As a consequence, studying individual protein isoforms is challenging, and inferred protein results are often abstracted to the gene-level or to groups of protein isoforms. Results:Here, we introduce IsoBayes, a novel statistical method to perform inference at the isoform level. Our method enhances the information available, by integrating mass spectrometry proteomics and transcriptomics data in a Bayesian probabilistic framework. To account for the uncertainty in the measurement process, we propose a two-layer latent variable approach: first, we sample if a peptide has been correctly detected (or, alternatively filter peptides); second, we allocate the abundance of such selected peptides across the protein(s) they are compatible with. This enables us, starting from peptide-level data, to recover protein-level data; in particular, we: i) infer the presence/absence of each protein isoform (via a posterior probability), ii) estimate its abundance (and credible interval), and iii) target isoforms where transcript and protein relative abundances significantly differ.We benchmarked our approach in simulations, and in two multi-protease real datasets: our method displays good sensitivity and specificity when detecting protein isoforms, its estimated abundances highly correlate with the ground truth, and can detect changes between protein and transcript relative abundances. Availability and implementation:IsoBayes is freely distributed as a Bioconductor R package, and is accompanied by an example usage vignette.
A critical part of the hepatitis B virus (HBV) life cycle is the packaging of the pregenomic RNA (pgRNA) into nucleocapsids. While this process is known to involve several viral elements, much less is known about the identities and roles of host proteins in this process. To better understand the role of host proteins, we isolated pgRNA and characterized its protein interactome in cells expressing either packaging-competent or packaging-incompetent HBV genomes. We identified over 250 host proteins preferentially associated with pgRNA from the packaging-competent version of the virus. These included proteins already known to support capsid formation, enhance viral gene expression, catalyze nucleocapsid dephosphorylation, and bind to the viral genome, demonstrating the ability of the approach to effectively reveal functionally significant host-virus interactors. Three of these host proteins, AURKA, YTHDF2, and ATR, were selected for follow-up analysis. RNA immunoprecipitation qPCR (RIP-qPCR) confirmed pgRNA-protein association in cells, and siRNA knockdown of the proteins showed decreased encapsidation efficiency. This study provides a template for the use of comparative RNA-protein interactome analysis in conjunction with virus engineering to reveal functionally significant host-virus interactions.
Introduction & Objective: High-risk single-nucleotide polymorphisms (SNPs) account for only ~20% of type 2 diabetes heritability, implying islet genes greatly influence one another to alter diabetes risk. Correlative analysis does not reveal conditional dependencies between key players regulating islet function whereas machine learning (ML) may do so. Methods: Using data from 374 genetically diverse mice, we derived models predicting islet function from protein abundance with gradient-boosted decision tree algorithms. Some mice (70%) were used to train the models and the rest were used to validate them. We also applied our models to similar data from other mouse studies and to data from humans to see if the models’ accuracies replicated. We then determined if the models included any known islet regulators, proteins with human orthologues possessing glycemia-related SNPs, and enriched for islet-relevant functional pathways. Finally, using the proteins’ predictive influences as meta-traits, we ran quantitative trait locus (QTL) scans to identify genomic regions altering the influence of proteins on islet function. Results: Our analysis revealed < 150 of the ~ 5000 detected proteins sufficiently modeled any one functional trait, with > 90% of a model’s prediction influenced by < 50 proteins. Some models predicted as well or better in the other mouse sets (R > 0.6) and reasonably well in human data (R > 0.5). ML analysis correctly identified the direction of effect for candidates (e.g. DPP8) where prior studies using correlation alone failed. Models for islet traits enriched for pathways potentially relevant for islet function and several unstudied proteins (e.g. CRELD2) present in multiple models have SNPs for glycemia-related traits. Finally, QTL scans revealed genomic regions that may alter the influence of key proteins on islet function and are far from those proteins’ genomic positions. Conclusions: ML can be used to identify conditional dependencies regulating islet function. Disclosure C. Emfinger: None. E. Freiberger: Employee; AbbVie Inc. M. Rabaglia: None. J. Kolic: None. S. Simonett: Employee; Exact Sciences. M. Shortreed: None. L. Clark: None. D. Stapleton: None. K. Schueler: None. K. Mitok: None. T.R. Price: None. J.J. Coon: Research Support; Agilent, Eli Lilly and Company, Merck & Co., Inc. L. Smith: None. M.J. Merrins: None. J.D. Johnson: None. M. Keller: None. A.D. Attie: None. Funding American Diabetes Association (7-21-PDF-157); National Institutes of Health (R01-DK101573-06); National Institutes of Health (GM070683); National Institutes of Health (P41-GM108538); National Institutes of Health (R35GM126914); National Institutes of Health (R35-GM118110); National Institutes of Health (T32-HL007936); National Institutes of Health (R01-DK113103); National Institutes of Health (R01DK127637); Canadian Institutes for Health Research (CIHR) operating grant 168857 CIHR Team Grant (ASD-179092/5-SRA-2021-1149-S-B); CIHR-JDRF Team (ASD-173663/5-SRA-2020-1059-S-B); CIHR Banting fellowship United States Department of Veterans Affairs Biomedical Laboratory Research and Development Service (I01BX005113)
Top-down proteomics, the characterization of intact proteoforms by tandem mass spectrometry, is the principal method for proteoform characterization in complex samples. Top-down proteomics relies on precursor isolation and subsequent gas-phase fragmentation to make proteoform identifications. While this strategy can produce highly detailed molecular information, the reliance on time-intensive tandem MS limits the speed with which proteoforms can be identified. We suggest that once proteoforms have been identified by top-down analysis in a system of interest, and archived in a system-specific Proteoform Atlas, subsequent analyses in that system can utilize the Atlas information to enable simpler and faster MS1-only identifications. We explore this idea here, using the E. coli ribosome as a model system of limited complexity. We used deep top-down analysis to construct an E. coli ribosomal Proteoform Atlas containing 2099 proteoforms from 52 of the 54 proteins that make up the E. coli ribosome. We show that using the Atlas enables confident MS1-only identifications of E. coli ribosomal proteoforms from E. coli that were perturbed by exposure to cold. Furthermore, this Atlas strategy identifies proteoforms up to 77% more rapidly compared to top-down identifications that require acquisition of both MS1 and MS2 spectra.
Identification of O-glycopeptides from tandem mass spectrometry data is complicated by the near complete dissociation of O-glycans from the peptide during collisional activation and by the combinatorial explosion of possible glycoforms when glycans are retained intact in electron-based activation. The recent O-Pair search method provides an elegant solution to these problems, using a collisional activation scan to identify the peptide sequence and total glycan mass, and a follow-up electron-based activation scan to localize the glycosite(s) using a graph-based algorithm in a reduced search space. Our previous O-glycoproteomics methods with MSFragger-Glyco allowed for extremely fast and sensitive identification of O-glycopeptides from collisional activation data but had limited support for site localization of glycans and quantification of glycopeptides. Here, we report an improved pipeline for O-glycoproteomics analysis that provides proteome-wide, site-specific, quantitative results by incorporating the O-Pair method as a module within FragPipe. In addition to improved search speed and sensitivity, we add flexible options for oxonium ion-based filtering of glycans and support for a variety of MS acquisition methods and provide a comparison between all software tools currently capable of O-glycosite localization in proteome-wide searches.
The identification of proteoforms by top-down proteomics requires both high quality fragmentation spectra and the neutral mass of the proteoform from which the fragments derive. Intact proteoform spectra can be highly complex and may include multiple overlapping proteoforms, as well as many isotopic peaks and charge states. The resulting lower signal-to-noise ratios for intact proteins complicates downstream analyses such as deconvolution. Averaging multiple scans is a common way to improve signal-to-noise, but mass spectrometry data contains artifacts unique to it that can degrade the quality of an averaged spectra. To overcome these limitations and increase signal-to-noise, we have implemented outlier rejection algorithms to remove outlier measurements efficiently and robustly in a set of MS1 scans prior to averaging. We have implemented averaging with rejection algorithms in the open-source, freely available, proteomics search engine MetaMorpheus. Herein, we report the application of the averaging with rejection algorithms to direct injection and online liquid chromatography mass spectrometry data. Averaging with rejection algorithms demonstrated a 45% increase in the number of proteoforms detected in Jurkat T cell lysate. We show that the increase is due to improved spectral quality, particularly in regions surrounding isotopic envelopes.