ALDH1A family enzymes (ALDH1A1, ALDH1A2, and ALDH1A3) catalyze retinoic acid synthesis, and their dysregulation is linked to disease. Selective inhibitors of these enzymes have been tested in drug discovery programs and one such compound, WIN18,446, was found to irreversibly inhibit ALDH1A2. WIN18,446 is a reversible male contraceptive in humans and animals. The inhibition of spermatogenesis by WIN18,446 is thought to be due to inhibition of ALDH1A. However, the mechanism of irreversible inhibition of ALDH1A2 by WIN18,446 is not known. A crystal structure obtained after incubating ALDH1A2 with WIN18,446 revealed a WIN18,446-derived metabolite covalently adducted to the catalytic cysteine, C320. Inspection of this structure suggested that the observed adduct is unstable and may be a metabolic intermediate stabilized under crystallographic conditions. In the current work, we tested this hypothesis. We identified and characterized an aldehyde metabolite of WIN18,446, which we designated M-54. M-54 is likely the metabolite of the intermediate observed in the crystal structure. Using a range of proteomics techniques, we identified a WIN18,446-derived ALDH1A2 protein adduct of mass 292.07 Da on C319 of ALDH1A2. This adduct may result from the reaction of the crystal structure metabolic intermediate. Using the identified mass, we probed human liver samples from multiple donors and found WIN18,446 specific adducts on cysteines within the ALDH1A1 and ALDH2 active site regions. The current study provides new insight into the metabolism of WIN18,446 and the mechanism of inhibition of ALDH1A2. We also demonstrate a proteomics workflow for identifying and validating drug-protein adducts of unknown mass.
Liquid chromatography-tandem mass spectrometry employing data-dependent acquisition (DDA) is a mature, widely used proteomics technique routinely applied to proteome profiling, protein-protein interaction studies, biomarker discovery, and protein modification analysis. Numerous tools exist for searching DDA data and myriad file formats are output as results. While some search and post processing tools include data visualization features to aid biological interpretation, they are often limited or tied to specific software pipelines. This restricts the accessibility, sharing and interpretation of data, and hinders comparison of results between different software pipelines. We developed Limelight, an easy-to-use, open-source, freely available tool that provides data sharing, analysis and visualization and is not tied to any specific software pipeline. Limelight is a data visualization tool specifically designed to provide access to the whole "data stack", from raw and annotated scan data to peptide-spectrum matches, quality control, peptides, proteins, and modifications. Limelight is designed from the ground up for sharing and collaboration and to support data from any DDA workflow. We provide tools to import data from many widely used open-mass and closed-mass search software workflows. Limelight helps maximize the utility of data by providing an easy-to-use interface for finding and interpreting data, all using the native scores from respective workflows.
Covalent protein adducts formed by drugs or their reactive metabolites are risk factors for adverse reactions, and inactivation of cytochrome P450 (CYP) enzymes. Characterization of drug-protein adducts is limited due to lack of methods identifying and quantifying covalent adducts in complex matrices. This study presents a workflow that combines data-dependent and data-independent acquisition (DDA and DIA) based liquid chromatography with tandem mass spectrometry (LC-MS/MS) to detect very low abundance adducts resulting from CYP mediated drug metabolism in human liver microsomes (HLMs). HLMs were incubated with raloxifene as a model compound and adducts were detected in 78 proteins, including CYP3A and CYP2C family enzymes. Experiments with recombinant CYP3A and CYP2C enzymes confirmed adduct formation in all CYPs tested, including CYPs not subject to time-dependent inhibition by raloxifene. These data suggest adducts can be benign. DIA analysis showed variable adduct abundance in many peptides between livers, but no concomitant decrease of unadducted peptides. This study sets a new standard for adduct detection in complex samples, offering insights into the human adductome resulting from reactive metabolite exposure. The methodology presented will aid mechanistic studies to identify, quantify and differentiate between adducts that result in adverse drug reactions and those that are benign.
Drugs are often metabolized to reactive intermedi-ates that form protein adducts. Adducts can inhibit protein activity,elicit immune responses, and cause life-threatening adverse drugreactions. The masses of reactive metabolites are frequentlyunknown, rendering traditional mass spectrometry-based proteo-mics approaches incapable of adduct identification. Here, wepresent Magnum, an open-mass search algorithm optimized foradduct identification, and Limelight, a web-based data processingpackage for analysis and visualization of data from all existingalgorithms. Limelight incorporates tools for sample comparisonsand xenobiotic-adduct discovery. We validate our tools with threedrug/protein combinations and apply our label-free workflow toidentify novel xenobiotic-protein adducts in CYP3A4. Our newmethods and software enable accurate identification of xenobiotic-protein adducts with no prior knowledge of adduct masses orprotein targets. Magnum outperforms existing label-free tools in xenobiotic-protein adduct discovery, while Limelight fulfills a majorneed in the rapidly developingfield of open-mass searching, which until now lacked comprehensive data visualization tools.
Proxl is an open-source web application for sharing, visualizing, and analyzing bottom-up protein cross-linking mass spectrometry data and results. Proxl's core features include comparing data sets, structural analysis, customizable and interactive data visualizations, access to all underlying mass spectrometry data, and quality-control tools. All features of Proxl are designed to be independent of specific cross-linker chemistry or software analysis pipelines. Proxl's sharing tools allow users to share their data with the public or securely restrict access to trusted collaborators. Since being published in 2016, Proxl has continued to be expanded and improved through active development and collaboration with cross-linking researchers. Some of Proxl's new features include a centralized, public site for sharing data, greatly expanded quality-control tools and visualizations, support for stable isotope-labeled peptides, and general improvements that make Proxl easier to use, data easier to share and import, and data visualizations more customizable. Source code and more information are found at http://proxl-ms.org/.
Metaproteomics is the characterization of all proteins being expressed by a community of organisms in a complex biological sample at a single point in time. Applications of metaproteomics range from the comparative analysis of environmental samples (such as ocean water and soil) to microbiome data from multicellular organisms (such as the human gut). Metaproteomics research is often focused on the quantitative functional makeup of the metaproteome and which organisms are making those proteins. That is: What are the functions of the currently expressed proteins? How much of the metaproteome is associated with those functions? And, which microorganisms are expressing the proteins that perform those functions? However, traditional protein-centric functional analysis is greatly complicated by the large size, redundancy, and lack of biological annotations for the protein sequences in the database used to search the data. To help address these issues, we have developed an algorithm and web application (dubbed "MetaGOmics") that automates the quantitative functional (using Gene Ontology) and taxonomic analysis of metaproteomics data and subsequent visualization of the results. MetaGOmics is designed to overcome the shortcomings of traditional proteomics analysis when used with metaproteomics data. It is easy to use, requires minimal input, and fully automates most steps of the analysis—including comparing the functional makeup between samples. MetaGOmics is freely available at https://www.yeastrc.org/metagomics/.
ProXL is a Web application and accompanying database designed for sharing, visualizing, and analyzing bottom-up protein cross-linking mass spectrometry data with an emphasis on structural analysis and quality control. ProXL is designed to be independent of any particular software pipeline. The import process is simplified by the use of the ProXL XML data format, which shields developers of data importers from the relative complexity of the relational database schema. The database and Web interfaces function equally well for any software pipeline and allow data from disparate pipelines to be merged and contrasted. ProXL includes robust public and private data sharing capabilities, including a project-based interface designed to ensure security and facilitate collaboration among multiple researchers. ProXL provides multiple interactive and highly dynamic data visualizations that facilitate structural-based analysis of the observed cross-links as well as quality control. ProXL is open-source, well-documented, and freely available at https://github.com/yeastrc/proxl-web-app .
Regulation of protein abundance is a critical aspect of cellular function, organism development, and aging. Alternative splicing may give rise to multiple possible proteoforms of gene products where the abundance of each proteoform is independently regulated. Understanding how the abundances of these distinct gene products change is essential to understanding the underlying mechanisms of many biological processes. Bottom-up proteomics mass spectrometry techniques may be used to estimate protein abundance indirectly by sequencing and quantifying peptides that are later mapped to proteins based on sequence. However, quantifying the abundance of distinct gene products is routinely confounded by peptides that map to multiple possible proteoforms. In this work, we describe a technique that may be used to help mitigate the effects of confounding ambiguous peptides and multiple proteoforms when quantifying proteins. We have applied this technique to visualize the distribution of distinct gene products for the whole proteome across 11 developmental stages of the model organism Caenorhabditis elegans. The result is a large multidimensional dataset for which web-based tools were developed for visualizing how translated gene products change during development and identifying possible proteoforms. The underlying instrument raw files and tandem mass spectra may also be downloaded. The data resource is freely available on the web at http://www.yeastrc.org/wormpes/.
Background Sequence feature annotations (e.g., protein domain boundaries, binding sites, and secondary structure predictions) are an essential part of biological research. Annotations are widely used by scientists during research and experimental design, and are frequently the result of biological studies. A generalized and simple means of disseminating and visualizing these data via the web would be of value to the research community. Findings Mason is a web site widget designed to visualize and compare annotated features of one or more nucleotide or protein sequence. Annotated features may be of virtually any type, ranging from annotating transcription binding sites or exons and introns in DNA to secondary structure or domain boundaries in proteins. Mason is simple to use and easy to integrate into web sites. Mason has a highly dynamic and configurable interface supporting multiple sets of annotations per sequence, overlapping regions, customization of interface and user-driven events (e.g., clicks and text to appear for tooltips). It is written purely in JavaScript and SVG, requiring no 3 rd party plugins or browser customization. Conclusions Mason is a solution for dissemination of sequence annotation data on the web. It is highly flexible, customizable, simple to use, and is designed to be easily integrated into web sites. Mason is open source and freely available at https://github.com/yeastrc/mason .
Accurate segregation of chromosomes during cell division is essential. The Dam1 complex binds kinetochores to microtubules and its oligomerization is required to form strong attachments. It is a key target of Aurora B kinase, which destabilizes erroneous attachments allowing subsequent correction. Understanding the roles and regulation of the Dam1 complex requires structural information. Here we apply cross-linking/mass spectrometry and structural modelling to determine the molecular architecture of the Dam1 complex. We find microtubule attachment is accompanied by substantial conformational changes, with direct binding mediated by the carboxy termini of Dam1p and Duo1p. Aurora B phosphorylation of Dam1p C terminus weakens direct interaction with the microtubule. Furthermore, the Dam1p amino terminus forms an interaction interface between Dam1 complexes, which is also disrupted by phosphorylation. Our results demonstrate that Aurora B inhibits both direct interaction with the microtubule and oligomerization of the Dam1 complex to drive error correction during mitosis.
The use of in vivo Förster resonance energy transfer (FRET) data to determine the molecular architecture of a protein complex in living cells is challenging due to data sparseness, sample heterogeneity, signal contributions from multiple donors and acceptors, unequal fluorophore brightness, photobleaching, flexibility of the linker connecting the fluorophore to the tagged protein, and spectral cross-talk. We addressed these challenges by using a Bayesian approach that produces the posterior probability of a model, given the input data. The posterior probability is defined as a function of the dependence of our FRET metric FRETR on a structure (forward model), a model of noise in the data, as well as prior information about the structure, relative populations of distinct states in the sample, forward model parameters, and data noise. The forward model was validated against kinetic Monte Carlo simulations and in vivo experimental data collected on nine systems of known structure. In addition, our Bayesian approach was validated by a benchmark of 16 protein complexes of known structure. Given the structures of each subunit of the complexes, models were computed from synthetic FRETR data with a distance root-mean-squared deviation error of 14 to 17 Å. The approach is implemented in the open-source Integrative Modeling Platform, allowing us to determine macromolecular structures through a combination of in vivo FRETR data and data from other sources, such as electron microscopy and chemical cross-linking.
BACKGROUND:As high throughput sequencing continues to grow more commonplace, the need to disseminate the resulting data via web applications continues to grow. Particularly, there is a need to disseminate multiple versions of related gene and protein sequences simultaneously--whether they represent alleles present in a single species, variations of the same gene among different strains, or homologs among separate species. Often this is accomplished by displaying all versions of the sequence at once in a manner that is not intuitive or space-efficient and does not facilitate human understanding of the data. Web-based applications needing to disseminate multiple versions of sequences would benefit from a drop-in module designed to effectively disseminate these data.FINDINGS:SnipViz is a client-side software tool designed to disseminate multiple versions of related gene and protein sequences on web sites. SnipViz has a space-efficient, interactive, and dynamic interface for navigating, analyzing and visualizing sequence data. It is written using standard World Wide Web technologies (HTML, Javascript, and CSS) and is compatible with most web browsers. SnipViz is designed as a modular client-side web component and may be incorporated into virtually any web site and be implemented without any programming.CONCLUSIONS:SnipViz is a drop-in client-side module for web sites designed to efficiently visualize and disseminate gene and protein sequences. SnipViz is open source and is freely available at https://github.com/yeastrc/snipviz.
To better understand the quantitative characteristics and structure of phenotypic diversity, we measured over 14,000 transcript, protein, metabolite, and morphological traits in 22 genetically diverse strains of Saccharomyces cerevisiae . More than 50% of all measured traits varied significantly across strains [false discovery rate (FDR) = 5%]. The structure of phenotypic correlations is complex, with 85% of all traits significantly correlated with at least one other phenotype (median = 6, maximum = 328). We show how high-dimensional molecular phenomics data sets can be leveraged to accurately predict phenotypic variation between strains, often with greater precision than afforded by DNA sequence information alone. These results provide new insights into the spectrum and structure of phenotypic diversity and the characteristics influencing the ability to accurately predict phenotypes.
Material Supplemental http://genome.cshlp.org/content/suppl/2013/06/25/gr.155762.113.DC1.html References http://genome.cshlp.org/content/23/9/1496.full.html#ref-list-1 This article cites 46 articles, 10 of which can be accessed free at: Open Access Open Access option. Genome Research Freely available online through the License Commons Creative . http://creativecommons.org/licenses/by-nc/3.0/ License (Attribution-NonCommercial 3.0 Unported), as described at , is available under a Creative Commons Genome Research This article, published in
Relocalization of proteins is a hallmark of the DNA damage response. We use high-throughput microscopic screening of the yeast GFP fusion collection to develop a systems-level view of protein reorganization following drug-induced DNA replication stress. Changes in protein localization and abundance reveal drug-specific patterns of functional enrichments. Classification of proteins by subcellular destination enables the identification of pathways that respond to replication stress. We analysed pairwise combinations of GFP fusions and gene deletion mutants to define and order two previously unknown DNA damage responses. In the first, Cmr1 forms subnuclear foci that are regulated by the histone deacetylase Hos2 and are distinct from the typical Rad52 repair foci. In a second example, we find that the checkpoint kinases Mec1/Tel1 and the translation regulator Asc1 regulate P-body formation. This method identifies response pathways that were not detected in genetic and protein interaction screens, and can be readily applied to any form of chemical or genetic stress to reveal cellular response pathways.
Laboratories engaged in computational biology or bioinformatics frequently need to run lengthy, multistep, and user-driven computational jobs. Each job can tie up a computer for a few minutes to several days, and many laboratories lack the expertise or resources to build and maintain a dedicated computer cluster.
Variation in RNA, protein, and metabolite levels among individuals is an important source of physiological and phenotypic differences within and between species. However, relatively little is known about the magnitude and genetic basis of these high-dimensional molecular phenotypes. Yeast provide an ideal model system for the genetic dissection of complex and quantitative traits, and whole-genome sequences are accumulating for dozens of Saccharomyces cerevisiae strains isolated from natural, industrial, and lab environments. We grew a diverse selection of sequenced strains in continuous culture and used a randomized and replicated study design. We exploited all the technologies in the Yeast Resource Center to obtain high quality and high coverage measurements of RNA, protein, metabolite, and morphological phenotypes. The resulting data sets provide a unique and powerful opportunity to combine comparative functional genomics data with comparative sequence analyses and delineate the genetic architecture of complex and quantitative phenotypes in yeast. Our initial analyses indicate that a high degree of strain-to-strain variation exists at all systems levels, and that this variation largely correlates with strain relatedness as measured by sequence comparison. Variation in RNA levels correlates with the corresponding peptides and related metabolites in complex ways. These experiments have resulted in an important large-scale data set of thousands of quantitative traits collected in a carefully designed randomized study, which will provide novel insights into the magnitude and patterns of natural variation of molecular and morphological phenotypes, as well as preliminary insights into their genetic basis.
Drugs are often metabolized to reactive intermediates that form protein adducts. Adducts can inhibit protein activity, elicit immune responses, and cause life-threatening adverse drug reactions. The masses of reactive metabolites are frequently unknown, rendering traditional mass spectrometry-based proteomics approaches incapable of adduct identification. Here, we present Magnum, an open-mass search algorithm optimized for adduct identification, and Limelight, a web-based data processing package for analysis and visualization of data from all existing algorithms. Limelight incorporates tools for sample comparisons and xenobiotic-adduct discovery. We validate our tools with three drug/protein combinations and apply our label-free workflow to identify novel xenobiotic-protein adducts in CYP3A4. Our new methods and software enable accurate identification of xenobiotic-protein adducts with no prior knowledge of adduct masses or protein targets. Magnum outperforms existing label-free tools in xenobiotic-protein adduct discovery, while Limelight fulfills a major need in the rapidly developing field of open-mass searching, which until now lacked comprehensive data visualization tools. H are constantly exposed to chemicals from their environment. Protein adducts result from covalent modification by xenobiotics, or their metabolites, and can also cause unintended toxicities and adverse drug reactions (ADRs). Adduct identification improves understanding of the mechanisms of ADRs and enables design of structural modifications to prevent them. Current identification methods, using radiolabeled compounds, trapping agents, immunological detection and mass spectrometry-based methods for detecting xenobiotic-protein adducts are labor-intensive and cannot sensitively measure the presence, abundance, and localization of protein adducts across a range of proteins without prior knowledge. Recent proteomics methods allow identification of unknown protein adducts but require isotopic labeling of test compounds or a complex combination of dataprocessing methods to reduce false discoveries and identify adducted peptides. Traditional database search algorithms search MS/MS spectra against protein sequences to identify the peptides and proteins that produced them. Known modifications can be identified if their masses are predefined. This is not possible for xenobiotic-protein adducts if their chemical composition is unknown. “Open-mass” search strategies solve this issue by allowing observed peptide masses to differ from identified peptide masses, returning mass differences as modifications of the identified peptides. This has allowed peptide spectrum matches (PSMs) to be made from a large proportion of previously unassigned spectra in shotgun proteomics data. However, few algorithms have been designed to identify xenobiotic-protein adducts that are typically low abundance and specific to the drug and treatment, and there are no graphical software tools to analyze and visualize data from open-mass searches. Here we present Magnum, a purpose-built xenobioticprotein adduct discovery algorithm, and Limelight, a webbased open modification analysis platform that rapidly highlights protein adducts that result from a specific treatment in a background of unrelated modifications. Limelight can combine and compare data from different pipelines empowering users to find the best tool or combination of tools for their specific application. We validate our new tools using three drug/protein combinations and compare results from Magnum to other open-mass search tools. We apply our workflow to identify novel xenobiotic-protein adducts in the P450 enzyme CYP3A4 resulting from exposure to raloxifene. Our software and Received: September 21, 2021 Accepted: February 10, 2022 Published: February 21, 2022 Article pubs.acs.org/ac © 2022 The Authors. Published by American Chemical Society 3501 https://doi.org/10.1021/acs.analchem.1c04101 Anal. Chem. 2022, 94, 3501−3509 D ow nl oa de d vi a 97 .1 13 .9 5. 10 1 on A pr il 1, 2 02 2 at 1 6: 08 :0 5 (U T C ). Se e ht tp s: //p ub s. ac s. or g/ sh ar in gg ui de lin es f or o pt io ns o n ho w to le gi tim at el y sh ar e pu bl is he d ar tic le s. workflow enable rapid and accurate identification of novel xenobiotic−protein adducts with no prior knowledge of adduct masses or protein targets. These tools provide a highly accelerated and statistically rigorous label-free workflow for the discovery and characterization of xenobiotic-protein adducts and can be incorporated into drug discovery pipelines and environmental toxicology screening. MATERIALS AND METHODS Reagents and Drug Incubations. The chemicals and reagents used in this study plus details of drug incubations are described in the Supplementary Methods section of the Supporting Information (SI). Sample Preparation. -lactam antibiotics and HSA: Aliquots (30 L) of control or drug treated HSA (0.5 mg/ mL) were reduced by adding 10 mM DTT, final concentration, and incubating at 37 °C for 30 min. Samples were alkylated with 16 mM iodoacetamide at room temp for 20 min in the dark. Tryptic digestion was done at 1:15 (enzyme:substrate) for 6 h at 37 °C in an Eppendorf Thermomixer with shaking (1000 rpm) prior to acidification with 250 mM HCl (final concentration). Samples were centrifuged at max speed in a benchtop microfuge for 10 min and supernatant transferred to autosampler vials and stored at −80 °C. Human plasma was prepared, treated, and digested as described in Supplementary Methods. Raloxifene and CYP3A4: Aliquots (100 L) of control or drug treated CYP3A4 incubation reaction mixture (63 g total protein) were prepared as described but digested for 4 h. A second set of digests were performed as described in Supplementary Methods and labeled “extra digest” in results. After digestion, solid phase extraction was done on all CYP3A4 samples using Oasis MCX cartridges (see Supplementary Methods for details). Mass Spectrometry. Sample digests (2 L ∼ 1 g) were loaded onto a 150 m Kasil fritted trap packed with 2 cm of ReprosilPur C18AQ (3 m bead diameter, Dr. Maisch) at a flow rate of 2 L per min. Separation used a self-packed 75 m i.d. 30 cm column. Peptides were eluted at 0.25 L/min using a standard or higher concentration (labeled “highB” in results) acetonitrile gradient. A QExactive HF or Exploris 480 (Thermo Fisher Scientific) was used to perform MS in data dependent mode. Data Processing. Acquired spectra were converted into mzML (for input to all algorithms except MODa) or mzXML (for input to MODa) using ProteoWizard’s msConvert. Proteins present in the samples were identified using Comet by standard closed searching against the entire human or E. coli proteomes, and smaller databases were made for subsequent open searching consisting only of proteins identified in initial comet searches by at least three peptides with a Percolator assigned q-value of ≤0.01. Decoy databases consisted of the corresponding set of reversed protein sequences and were provided to algorithms requiring pregenerated decoy sequences. All data are filtered at a false discovery rate (FDR) of 1% unless otherwise stated. Detailed procedures and parameters for each algorithm can be found in Supplementary Methods. Native search results from all pipelines were converted to Limelight XML prior to uploading to the Limelight web application. We have written converters for many pipelines including Comet, Percolator, the Trans-Proteomic Pipeline (TPP), Crux, MSFragger, open-pFind, Comet-PTM, MetaMorpheus, MODa, TagGraph, and Magnum (th i s pape r) . A cur ren t l i s t o f conve r t e r s i s a v a i l a b l e o n o u r d o c u m e n t a t i o n W e b s i t e : https://limelight-ms.readthedocs.io/. Further details can be found in Supplementary Note 3. Quanti cation of CYP3A4 and P450-Reductase Adducts. Peptides were quantified using Skyline as previously described. Full details can be found in Supplementary Methods, processed data plus Limelight links can be found in Supplementary File 2, Sheet 3, and a complete interactive Skyline session is available on Panorama (https:// panoramaweb.org/CYP3A4-raloxifene.url). Software Availability. Magnum is written in C++ and is open source and freely available at http://magnum-ms.org. Full source code plus precompiled Magnum binaries are available for Windows and Linux. Docker images for the Magnum/Percolator/Limelight pipeline and associated documentation can be found at https://limelight-ms.readthedocs. io/en/latest/tutorials/magnum-pipeline.html. Magnum outputs results either as simple tab-delimited text or PepXML format for potential use with existing software supporting this format. For further details of Magnum see Supplementary Note 1. Limelight is written in Java and TypeScript and is open source and freely available at https://limelight-ms.org/. Limelight source code plus preconfigured Docker containers are provided for running Limelight and Limelight XML converters. Extensive documentation is at https://limelightms.readthedocs.io/. Users not able to run their own Limelight installation may use a generally available installation at https://use.limelight-ms.org/. For full details of Limelight see Supplementary Note 3. Data Availability. All raw and processed data discussed in this paper are available via Limelight at https://limelight. yeastrc.org/limelight/p/adduct-discovery. In addition, complete search algorithm configuration files, fasta search databases, raw search output, and raw MS data files were deposited to the ProteomeXchange Consortium via the PRIDE partner repository with the data set identifier PXD025019. Full Skyline quantification of CYP3A4/raloxifene was deposited to the ProteomeXchange Consortium via Panorama Public and is available at https://panoramaweb.org/CYP3A4-raloxifene.url with the data set identifier PXD024932.