ALDH1A family enzymes (ALDH1A1, ALDH1A2, and ALDH1A3) catalyze retinoic acid synthesis, and their dysregulation is linked to disease. Selective inhibitors of these enzymes have been tested in drug discovery programs and one such compound, WIN18,446, was found to irreversibly inhibit ALDH1A2. WIN18,446 is a reversible male contraceptive in humans and animals. The inhibition of spermatogenesis by WIN18,446 is thought to be due to inhibition of ALDH1A. However, the mechanism of irreversible inhibition of ALDH1A2 by WIN18,446 is not known. A crystal structure obtained after incubating ALDH1A2 with WIN18,446 revealed a WIN18,446-derived metabolite covalently adducted to the catalytic cysteine, C320. Inspection of this structure suggested that the observed adduct is unstable and may be a metabolic intermediate stabilized under crystallographic conditions. In the current work, we tested this hypothesis. We identified and characterized an aldehyde metabolite of WIN18,446, which we designated M-54. M-54 is likely the metabolite of the intermediate observed in the crystal structure. Using a range of proteomics techniques, we identified a WIN18,446-derived ALDH1A2 protein adduct of mass 292.07 Da on C319 of ALDH1A2. This adduct may result from the reaction of the crystal structure metabolic intermediate. Using the identified mass, we probed human liver samples from multiple donors and found WIN18,446 specific adducts on cysteines within the ALDH1A1 and ALDH2 active site regions. The current study provides new insight into the metabolism of WIN18,446 and the mechanism of inhibition of ALDH1A2. We also demonstrate a proteomics workflow for identifying and validating drug-protein adducts of unknown mass.
Aging reshapes the cellular and molecular landscape of mammalian tissues. These changes can be progressive, preceding linearly with age, or occur as abrupt transitions of the course of lifespan. To investigate the age-dependent cellular and molecular shifts we profiled matched proteomes and transcriptomes from male and female murine spleens across eight time points, from stable adults through late life. The spleen was chosen to integrate understanding of age-dependent changes associated with immune surveillance, inflammaging, and immune-related proteostasis. Male and female mice follow distinct aging trajectories particularly in protein-RNA correlation in late life, reflecting both compositional shifts and failure of post-transcriptional buffering. To investigate whether these changes could be attributed to specific cell-types within the spleen, we developed Celestial, a machine-learning framework to identify cell-type-specific changes in bulk tissue samples. We found that age-related bulk molecular changes could be attributed in part to compositional remodeling of cell-types-expansion of GZMK+ CD8+ T cells and C1Q+ macrophages alongside naive T cell and global B cell loss. These results demonstrate that cell-type-aware interpretation can inform bulk multi-omic data for accurate mechanistic inference in heterogeneous tissues undergoing complex molecular remodeling.
Cells release membrane-bound extracellular vesicles into the bloodstream laden with proteins that may reflect their physiological state. How this circulating EV proteome changes across life remains poorly understood. Identifying molecular signatures of aging in accessible biofluids could facilitate earlier intervention and monitoring of age-related disease. Many circulating aging proteome studies rely on affinity-based platforms which suffer from poor cross-species translation, ambiguous signal attribution, and inconsistent agreement between platforms. Here, we present a characterization of the aging plasma EV proteome from a cross-sectional cohort of 86 male and female C57BL/6J mice (5-31 months). We leveraged a species-agnostic EV enrichment (Mag-Net) and mass spectrometry to detect 2,575 protein groups from 15,969 peptides. Protein abundance heterogeneity increased with age and the abundance of 272 proteins were significantly correlated with chronological age including established senescence and frailty markers. Proteins increasing with age were enriched in genome maintenance pathways, while those decreasing were associated with the extracellular matrix organization and lipid metabolism. Notably, several of the strongest age-increased proteins converged on Alzheimer's and Parkinson's disease pathology. We observed sexual divergence in the aging EV proteome not previously characterized at this resolution. A proteomic clock built from this data accurately predicts chronological age, and peptide-level analysis reveals aging signals invisible at protein-level. These findings demonstrate that EV-enriched plasma proteomics can identify known aging markers, reveal novel sex-specific age-related changes, and generate predictive models of chronological age. This study provides a species-agnostic foundation for proteomic clocks that complement epigenetic approaches to monitor aging and evaluate healthspan.
Obesity is a major public health challenge affecting an ever-increasing proportion of the global population. It is associated with numerous comorbidities. Progressive expansion and remodeling of adipose tissue may lead to depot specific changes in adipose tissue biology and energy partitioning. Such changes likely precede the development of obesity-related complications. To facilitate a deeper understanding of adipose tissue biology, a comprehensive quantitative proteomic dataset is presented at the peptide and protein level. Data-independent acquisition LC-MS/MS data were acquired from matched subcutaneous and omental adipose tissues from metabolically healthy individuals with no comorbidities and covering a wide range of body mass indexes. Adipose tissue samples were collected during elective surgeries and immediately processed for histology or frozen until proteomic analysis. Internal and external quality control systems ensured high quality data. All data presented are available via ProteomeXchange. This dataset will allow new insights into biological changes that evolve with increasing adiposity captured before the onset of comorbidities. Matched sampling across fat depots provides an opportunity to uncover depot-specific physiological signatures.
Dogma suggests protein quantification is a pre-requisite to LC-MS/MS based proteomics studies. Such quantification allows a standardized ratio of sample to digestion enzyme and enables physical normalization of protein digest loaded onto the mass spectrometer for analysis. Most proteomics studies include these steps. However, there are significant costs in time, money and experimental complexity, associated with performing protein quantification and physical normalization for every sample, especially for larger studies. Proteomics data analysis pipelines typically include computational normalization strategies to compensate for unavoidable systematic biases. These strategies also have the potential to compensate for avoidable variation such as omitting sample amount normalization. Here we investigate the effects of either physically normalizing the amount of protein for each individual sample or leaving it unnormalized. Our results show the relationship between increased protein amount variation in sample input, and the variance of quantified relative abundances of peptides and proteins output after data analysis. The experiments presented here suggest that protein quantification and physical normalization steps can be omitted from some quantitative proteomic experiments without incurring an unacceptable increase in measurement variability after computational normalization has been applied. This work will enable important time and cost saving optimizations to be made to many proteomics workflows.
Abstract Background The developmental pattern of the crystalline lens provides a unique model to study biological aging and its effects on the posttranslational modification of long-lived proteins. The orderly differentiation of lens fiber cells leads to a spatiotemporal gradient where mature, organelle-free fiber cells are packed in the lens nucleus surrounded by con-centric rings of successively younger fiber cells in the cortex. Methods Pig lenses were separated into six layers by dissolution in a hypotonic buffer. The changes in protein abundance, oxidation, and phosphorylation that occur across the spatiotemporal gradient of the lens were assessed quantitatively by using data independent acquisition label-free proteomic analysis of these six fractions. Results Expected changes in protein abundance of major lens protein which reflect the maturation process of lens fiber cells across the spatiotemporal gradient were found. Significant differences were noted in phosphorylation sites on crystallins, phakinin, and actin. Significant changes in oxidation of residues on essential lens proteins, as well as several glycolytic enzymes, were found across the spatiotemporal gradient. Conclusion Dissolution of the lens followed by high resolution data-independent acquisition proteomics is a powerful technique for spatial mapping of protein abundance and posttranslational modification changes in the lens. The oxidation and phosphorylation sites noted in this study may play important roles in both lens development and cataractogenesis.
De novo peptide sequencing detects peptides directly from tandem mass spectra without a protein sequence database, and deep learning has substantially advanced its performance. Casanovo, one such widely used model, is distributed as a Python command-line program. Consequently, installation, GPU and dependency configuration, and manual parameterization can be challenging for many bench scientists and are a recurring source of errors. Interpreting and validating the resulting predictions poses a further challenge. We present CasanovoGUI, an open-source Java-based desktop application that makes all of Casanovo's main analysis functions available through a point-and-click interface on Windows, macOS, and Linux. On first use, CasanovoGUI automatically installs a private Python environment and Casanovo with a GPU-matched build, requiring no prior software setup. The GUI provides access to Casanovo's analysis functions and configuration parameters, streams live progress, and integrates results interpretation: annotated spectra with per-residue confidence scores in the PDV viewer, and mismatch-tolerant mapping of de novo peptides back to a reference proteome. CasanovoGUI is available at https://github.com/Noble-Lab/CasanovoGUI.
Ionizing radiation induces molecular responses that may be used to estimate exposure when physical dosimeters are unavailable. Here we present two large-scale proteomics datasets generated from mouse dorsal skin punch samples collected following controlled X-ray exposures spanning multiple doses, dose rates, and post-exposure time points. Experiment 1 comprised 96 samples (including 16 reference samples) collected 6 days after exposure to 0-75 cGy delivered at either 30 or 300 cGy/min. Experiment 2 comprised 936 samples (including 236 reference samples) exposed to 0-100 cGy at either 3 or 28 cGy/min dose rates and harvested between 7 and 150 days post-exposure. All samples were processed using a standardized workflow involving automated bead-based digestion and data-independent acquisition mass spectrometry. The datasets include multiple pooled reference sample types, process controls, and system suitability standards ensuring high quality data. All data presented are available via ProteomeXchange at several levels of processing, from raw files through normalized peptide- and protein-level abundance matrices suitable for biomarker discovery and machine learning applications. This dataset will facilitate generation of new insights into the biological changes and molecular signatures resulting from X-ray exposure in mice and may also help inform future studies in humans.
The traditional approaches to handling missing values in DIA proteomics are to either remove high-missingness proteins or impute them with statistical procedures. Both have their disadvantages-removal can limit statistical power, while imputation can introduce spurious correlations or dilute signal. We present an alternative approach based on imputing peptide retention times (RTs) rather than quantitations. For each missing value, we impute the RT boundaries, then obtain a quantitation by integrating the chromatographic signal within the imputed boundaries. Our method yields more accurate quantitations than existing proteomics imputation methods. RT boundary imputation also identifies differentially abundant peptides from key Alzheimer's genes that were not identified with library search alone. RT boundary imputation improves the ability to estimate radiation exposure in biological tissues. RT boundary imputation significantly increases the number of peptides with quantitations, leading to increases in statistical power. Finally, RT boundary imputation better quantifies low abundance peptides than library search alone. Our RT boundary imputation method, called Nettle, is available as a standalone tool.
E3 ubiquitin ligases play a crucial role in modulating receptor stability and signaling at the cell surface, yet the mechanisms governing their substrate specificity remain incompletely understood. Mahogunin Ring Finger 1 (MGRN1) is a membrane-tethered E3 ligase that fine-tunes signaling sensitivity by targeting surface receptors for ubiquitination and degradation. Unlike cytosolic E3 ligases, membrane-tethered E3s require transmembrane adapters to selectively recognize and regulate surface receptors, yet few such ligases have been studied in detail. While MGRN1 is known to regulate the receptor Smoothened (SMO) within the Hedgehog pathway through its interaction with the transmembrane adapter Multiple Epidermal Growth Factor-like 8 (MEGF8), the broader scope of its regulatory network has been speculative. Here, we identify Attractin (ATRN) and Attractin-like 1 (ATRNL1) as additional transmembrane adapters that recruit MGRN1 and regulate cell surface receptor turnover. Through co-immunoprecipitation, we show that ATRN and ATRNL1 likely interact with the RING domain of MGRN1. Functional assays reveal that MGRN1 requires these transmembrane adapters to ubiquitinate and degrade the melanocortin receptors MC1R and MC4R, in a process analogous to its regulation of SMO. Loss of MGRN1 leads to increased surface and ciliary localization of MC4R in fibroblasts and elevated MC1R levels in melanocytes, with the latter resulting in enhanced eumelanin production. These findings expand the repertoire of MGRN1-regulated receptors and provide new insight into a shared mechanism by which membrane-tethered E3 ligases utilize transmembrane adapters to dictate substrate receptor specificity. By elucidating how MGRN1 selectively engages with surface receptors, this work establishes a broader framework for understanding how this unique class of E3 ligases fine-tunes receptor homeostasis and signaling output.
Abstract In May and June of 2021, marine microbial samples were collected for DNA sequencing in East Sound, WA, USA every 4 hours for 22 days. This high temporal resolution sampling effort captured the last 3 days of a Rhizosolenia sp. bloom, the initiation and complete bloom cycle of Chaetoceros socialis (8 days), and the following bacterial bloom (2 days). Metagenomes were completed on the time series, and the dataset includes 128 size-fractionated microbial samples (0.22–1.2 µm), providing gene abundances for the dominant members of bacteria, archaea, and viruses. This dataset also has time-matched nutrient analyses, flow cytometry data, and physical parameters of the environment at a single point of sampling within a coastal ecosystem that experiences regular bloom events, facilitating a range of modeling efforts that can be leveraged to understand microbial community structure and their influences on the growth, maintenance, and senescence of phytoplankton blooms.
Casanovo is a state-of-the-art deep learning model for de novo peptide sequencing from mass spectrometry and proteomics data. Here, we report on a series of enhancements to Casanovo, aimed at improving the interpretability of the scores assigned to predicted peptides, generalizing the software for use in database searches, speeding up training and prediction runtimes, and providing workflows and visualization tools to facilitate adoption of Casanovo and interpretation of its results. Our goal is to make Casanovo accurate and easy to use for applications such as metaproteomics, antibody sequencing, immunopeptidomics, and the discovery of novel peptide sequences in standard proteomics analyses. Casanovo is available as open source at https://github.com/Noble-Lab/casanovo.
A pressing statistical challenge in the field of mass spectrometry proteomics is how to assess whether a given software tool provides accurate error control. Each software tool for searching such data uses its own internally implemented methodology for reporting and controlling the error. Many of these software tools are closed source, with incompletely documented methodology, and the strategies for validating the error are inconsistent across tools. In this work, we identify three different methods for validating false discovery rate (FDR) control in use in the field, one of which is invalid, one of which can only provide a lower bound rather than an upper bound, and one of which is valid but under-powered. The result is that the field has a very poor understanding of how well we are doing with respect to FDR control, particularly for the analysis of data-independent acquisition (DIA) data. We therefore propose a theoretical formulation of entrapment experiments that allows us to rigorously characterize the behavior of the various entrapment methods. We also propose a more powerful method for evaluating FDR control, and we employ that method, along with other existing techniques, to characterize a variety of popular search tools. We empirically validate our entrapment analysis in the fairly well-understood DDA setup before applying it in the DIA setup. We find that none of the DIA search tools consistently controls the FDR at the peptide level, and the tools struggle particularly with analysis of single cell datasets.
A critical challenge in mass spectrometry proteomics is accurately assessing error control, especially given that software tools employ distinct methods for reporting errors. Many tools are closed-source and poorly documented, leading to inconsistent validation strategies. Here we identify three prevalent methods for validating false discovery rate (FDR) control: one invalid, one providing only a lower bound, and one valid but under-powered. The result is that the proteomics community has limited insight into actual FDR control effectiveness, especially for data-independent acquisition (DIA) analyses. We propose a theoretical framework for entrapment experiments, allowing us to rigorously characterize different approaches. Moreover, we introduce a more powerful evaluation method and apply it alongside existing techniques to assess existing tools. We first validate our analysis in the better-understood data-dependent acquisition setup, and then, we analyze DIA data, where we find that no DIA search tool consistently controls the FDR, with particularly poor performance on single-cell datasets.
Data-independent acquisition (DIA)-based mass spectrometry is becoming an increasingly popular mass spectrometry acquisition strategy for carrying out quantitative proteomics experiments. Most of the popular DIA search engines make use of in silico generated spectral libraries. However, the generation of high-quality spectral libraries for DIA data analysis remains a challenge, particularly because most such libraries are generated directly from data-dependent acquisition (DDA) data or are from in silico prediction using models trained on DDA data. In this study, we introduce Carafe, a tool that generates high-quality experiment-specific in silico spectral libraries by training deep learning models directly on DIA data. We demonstrate the performance of Carafe on a wide range of DIA datasets, where we observe improved fragment ion intensity prediction and peptide detection relative to existing pretrained DDA models. To make Carafe more accessible to the community, we integrate Carafe into the widely used Skyline tool.
Previously, we reconstituted a minimal functional kinetochore from recombinant Saccharomyces cerevisiae proteins that was capable of transmitting force from dynamic microtubules to nucleosomes containing the centromere-specific histone variant Cse4 (Hamilton et al. 2020). This work revealed two paths of force transmission through the inner kinetochore: through Mif2 and through the Okp1/Ame1 complex (OA). Here, using a chimeric DNA sequence that contains crucial centromere-determining elements of the budding yeast point centromere, we demonstrate that the presence of centromeric DNA sequences in Cse4-containing nucleosomes significantly strengthens OA-mediated linkages. Our findings indicate that centromeric sequences are important for the transmission of microtubule-based forces to the chromosome.
Extracellular vesicles (EVs) in plasma are composed of exosomes, microvesicles, and apoptotic bodies. We report a plasma EV enrichment strategy using magnetic beads called Mag-Net. Proteomic interrogation of this plasma EV fraction enables the detection of proteins that are beyond the dynamic range of liquid chromatography-mass spectrometry of unfractionated plasma. Mag-Net is robust, reproducible, inexpensive, and requires <100 μL plasma input. Coupled to data-independent mass spectrometry, we demonstrate the measurement of >37,000 peptides from >4,000 proteins. Using Mag-Net on a pilot cohort of patients with neurodegenerative disease and healthy controls, we find 204 proteins that differentiate (q-value < 0.05) patients with Alzheimer's disease dementia (ADD) from those without ADD. There are also 310 proteins that differ between individuals with Parkinson's disease and without. Using machine learning we distinguish between individuals with ADD and not ADD with an area under the receiver operating characteristic curve (AUROC) = 0.98 ± 0.06.
A core computational challenge in the analysis of mass spectrometry data is the de novo sequencing problem, in which the generating amino acid sequence is inferred directly from an observed fragmentation spectrum without the use of a sequence database. Recently, deep learning models have made substantial advances in de novo sequencing by learning from massive datasets of high-confidence labeled mass spectra. However, these methods are designed primarily for data-dependent acquisition experiments. Over the past decade, the field of mass spectrometry has been moving toward using data-independent acquisition (DIA) protocols for the analysis of complex proteomic samples owing to their superior specificity and reproducibility. Hence, we present a de novo sequencing model called Cascadia, which uses a transformer architecture to handle the more complex data generated by DIA protocols. In comparisons with existing approaches for de novo sequencing of DIA data, Cascadia achieves substantially improved performance across a range of instruments and experimental protocols.
Mahogunin ring finger 1 (MGRN1) is a membrane-tethered E3 ligase that fine-tunes signaling sensitivity by targeting surface receptors for ubiquitylation and degradation. Although MGRN1 is known to regulate the Hedgehog signaling effector Smoothened (SMO) via the transmembrane adapter multiple epidermal growth factor-like 8 (MEGF8), the broader scope of its regulatory network has been speculative. Here, we identify attractin (ATRN) and attractin-like 1 (ATRNL1) as additional transmembrane adapters that recruit MGRN1 and regulate cell surface receptor turnover. Through coimmunoprecipitation, we show that ATRN interacts with the RING domain of MGRN1. Functional assays suggest that ATRN and ATRNL1 work with MGRN1 to promote the ubiquitylation and degradation of the melanocortin receptors MC1R and MC4R, in a process analogous to its regulation of SMO. Loss of MGRN1 or ATRN leads to increased surface and ciliary localization of MC4R in fibroblasts and elevated MC1R levels in melanocytes, resulting in enhanced eumelanin production. These findings expand the known repertoire of MGRN1-regulated receptors and provide new insight into a shared mechanism by which membrane-tethered E3 ligases utilize transmembrane adapters to facilitate substrate receptor specificity.