The proteome is a dynamic landscape of proteoforms arising from genetic mutations, alternative splicing, and post-translational modifications (PTMs), which collectively drive biological function and disease phenotypes. Mass spectrometry (MS)-based proteomics has emerged as an essential technique for elucidating this molecular complexity. Although bottom-up proteomics enables deep protein identification and quantification through peptide-level analysis, it disrupts molecular connectivity and introduces a peptide-to-protein inference problem, which is suboptimal for proteoform analysis. Top-down proteomics (TDP) offers a complementary approach by analyzing intact proteins, preserving molecular connectivity, and enabling direct characterization and quantification of proteoforms. This capability is increasingly vital for understanding heterogeneous human diseases. Here, we review the evolving role of TDP in biomedical research, highlighting studies that revealed proteoform-level alterations, identified candidate biomarkers, and advanced our understanding of the roles of proteoforms in human diseases.
Abstract Spirochaete pathogens are among the most invasive bacteria known, causing syphilis, Lyme disease, and leptospirosis. Their tissue penetration depends on periplasmic flagellar filaments that, unlike other bacterial flagella, are encased in a spirochaete-specific multi-protein sheath and deform the cell body into motile waves. How these filaments achieve the mechanical properties needed for invasive motility has remained unclear. Here we determine complete atomic structures of the Leptospira endoflagellar filament, revealing an elaborate sheath of 9 to 12 distinct asymmetrically arranged proteins. We show that the flagellin variant forming the filament core determines sheath composition, producing curvatures ranging from ~3.5 µm −1 to ~5.6 µm −1 . The lower-curvature architecture, employed by pathogenic Leptospira interrogans , proves essential for motility in viscous environments and during infection. Thus, Leptospira achieves environment-specific motility through modular core–sheath coupling, linking atomic-scale structural plasticity to large-scale changes in swimming behaviour. Conservation of key sheath components suggests this mechanism may extend across spirochaetes.
Ribosome hibernation preserves translation machinery during stress, yet its mechanisms in Archaea remain poorly defined. Using cryo-EM analysis, we studied hibernation pathways in Pyrococcus abyssi stressed cells. We identified HibA, a previously unrecognized family of hibernation factors widespread in Archaea. HibA consists of a bacterial-like HPF/RaiA domain fused to a Cystathionine Beta Synthase module. Unexpectedly, HibA binds to the ribosome in three different conformations, occupying the A, P and E sites of tRNAs, as well as that of mRNA, enhancing its ability to protect the ribosome from degradation. Idle ribosomes also frequently accumulate the archaeal homolog of eukaryotic ribosome maturation protein SBDS (aSBDS), suggesting that stressed archaeal cells may engage parallel hibernation routes in which aSBDS can complement HibA. Deletion of hibA in Thermococcus barophilus delays recovery from stationary phase and reduces 70S ribosome pools, establishing its role in ribosome preservation. Taxonomic profiling shows that many archaeal lineages encode distinct repertoires of ribosome-associated protection factors, underscoring the modular and multi-layered nature of archaeal hibernation systems. In addition, a comprehensive phylogenetic analysis highlights the evolutionary relationships between prevalent ribosome hibernation factors across Bacteria and Archaea.
Therapeutic antibodies, primarily immunoglobulin G-based monoclonal antibodies, are developed to treat cancer, autoimmune disorders, and infectious diseases. Their large size, structural complexity, and heterogeneity pose significant analytical challenges, requiring advanced characterization techniques. This review traces the 30-year evolution of top-down (TD) and middle-down (MD) mass spectrometry (MS) for antibody analysis, beginning with their initial applications and highlighting key advances and challenges throughout this period. TD MS allows for the analysis of intact antibodies, and MD MS performs analysis of the antibody subunits, even in complex biological samples. Both approaches preserve critical quality attributes such as sequence integrity, post-translational modifications (PTMs), disulfide bonds, and glycosylation patterns. Key milestones in TD and MD MS of antibodies include the use of structure-specific enzymes for subunit generation, the implementation of high-resolution mass spectrometers, and the adoption of non-ergodic ion activation methods such as electron transfer dissociation (ETD), electron capture dissociation (ECD), ultraviolet photodissociation (UVPD), and matrix-assisted laser desorption/ionization in-source decay (MALDI-ISD). The combination of complementary dissociation methods and consecutive ion activation approaches has further enhanced TD/MD MS performance. The current TD MS record of antibody sequencing with terminal product ions is about 60% sequence coverage obtained using the activated ion-ETD approach on a high-resolution MS platform. Current MD MS analyses with about 95% sequence coverage were achieved using combinations of ion activation and dissociation techniques. The review explores TD and MD MS analysis of novel mAb modalities, including antibody-drug conjugates, bispecific antibodies, endogenous antibodies from biofluids, and immunoglobulin A and M-type classes.
The archaeal ribosome is of the eukaryotic type. TACK and Asgard superphyla, the closest relatives of eukaryotes, have ribosomes containing eukaryotic ribosomal proteins not found in other archaea, eS25, eS26 and eS30. Here, we investigate the case of Saccharolobus solfataricus, a TACK crenarchaeon, using mainly leaderless mRNAs. We characterize the small ribosomal subunit of S. solfataricus bound to SD-leadered or leaderless mRNAs. Cryo-EM structures show eS25, eS26 and eS30 bound to the small subunit. We identify two ribosomal proteins, aS33 and aS34, and an additional domain of eS6. Leaderless mRNAs are bound to the small subunit with contribution of their 5'-triphosphate group. Archaeal eS26 binds to the mRNA exit channel wrapped around the 3' end of rRNA, as in eukaryotes. Its position is not compatible with an SD:antiSD duplex. Our results suggest a positive role of eS26 in leaderless mRNAs translation and possible evolutionary routes from archaeal to eukaryotic translation.
Generating top-down tandem mass spectra (MS/MS) for complex mixtures of proteoforms has become possible through improvements in fractionation, on-line separation, dissociation, and mass analysis. The algorithms to match MS/MS to sequences have undergone a parallel evolution, with both spectral alignment and peak matching being paired with diverse methods for scoring proteoform-spectral matches (PrSMs). This study assesses state-of-the-art algorithms for top-down identification through three distinct challenges. The first is identifying a large yield of PrSMs while controlling false discovery rate (FDR) in identifying thousands of proteoforms from complex cell lysates via four software workflows: ProSight Proteome Discoverer, TopPIC, Informed Proteomics, and pTop. The second is the deconvolution of data from both Thermo Orbitrap-class and Bruker maXis Q-TOF instruments to produce consistent precursor charge and mass determinations while generating fragment mass lists to optimize identification. The third attempts to detect diverse post-translational modifications (PTMs) in proteoforms from bovine milk and human ovarian tissue. The data demonstrate that existing software suites produce admirable sensitivity, in some cases identifying a third of collected MS/MS with FDR controlled below 2%; the overlap in these PrSMs, however, illustrates real value in searching data with multiple search engines. Differences among identification workflows seem to result from each search algorithm incorporating its own deconvolution algorithm. By transmitting deconvolution data from multiple deconvolution routes (Thermo Xtract, Bruker Auto MSn, Mascot Distiller, TopFD, and FLASHDeconv) to the downstream TopPIC search algorithm, we were able to detect common causes of deconvolution disagreement. The detection of PTMs was very inconsistent among search algorithms, with some workflows suggesting as little as 1% of PrSMs from bovine milk were singly-phosphorylated while other workflows found that 18% of PrSMs were singly-phosphorylated. Taken together, these results make a strong argument for top-down researchers to adopt a standard practice of analyzing each MS/MS experiment with at least two different search engines.
Tubulin polyglutamylation is a key feature of eukaryotic cilia and flagella that is essential for their function. The diversity of enzymes catalyzing polyglutamylation with different specificities inspired the hypothesis of the tubulin code. In the protist parasite Trypanosoma brucei, nine different glutamylase enzymes are potentially involved in tubulin glutamylation. To decipher the trypanosome tubulin code generated by this diversity, we aimed at determining tubulin glutamylation patterns by robust mass spectrometry (MS)-based proteomics. MS approaches exist for many post-translational modifications but none for the chemically complex polyglutamylation. We therefore optimized a nanoLC-MS/MS pipeline from sample preparation to data analysis using synthetic peptides for quantification. Our approach enabled the quantification of C-terminal tubulin peptides with up to 11 supplementary glutamates on α-, and five on β-tubulin from the flagellum of T. brucei. In addition to the known E445 on α- and E435 on β-, a novel glutamylation site of β-tubulin was discovered at E438. Furthermore, our data revealed an increase in enzymatic detyrosination with increasing length of the glutamate chains, especially for α-tubulin. This indicates cross-talk between the modifications and different detyrosination rates of the two tubulin types. Our efficient analytical pipeline advances understanding of the tubulin code in T. brucei.
The Omnitrap-Orbitrap-Booster (OOB) mass spectrometry (MS) platform was developed to advance the top-down (TD) MS analysis of proteins. It integrates a multimodal tandem mass spectrometry (MS/MS) ion trap system (Omnitrap), a high-resolution Orbitrap Fourier transform mass spectrometer (FTMS), and a high-performance data acquisition system (FTMS Booster) to improve fragmentation efficiency and spectral quality by increasing the signal-to-noise (S/N) ratio of product ions. In this study, we evaluate the OOB platform for the electron capture dissociation (ECD)-based TD MS analysis of a P15 multiple myeloma antibody light chain. Single precursor charge state analysis of P15 23+ yielded relatively high sequence coverage of 68%, albeit indicating a limitation caused by the overlap of certain product ions with the charge reduced precursors. The corresponding method development, leveraging consecutive analysis of multiple precursor charge states (15+ to 19+) across triplicate LC-MS/MS runs on the OOB platform, enhanced P15 sequence coverage to 93%, demonstrating its capacity for comprehensive protein characterization. In addition, we demonstrate that the obtained ECD-based TD MS performance on the OOB platform for P15 light chain is comparable to the "gold-standard" electron transfer/higher-energy collision dissociation (EThcD)-based TD MS on an Orbitrap Eclipse. Serendipitously, ECD exhibits a lower spectral peak density (i.e., reduced spectral congestion) due to reduced redundancy of product ions. These results establish the OOB platform as a powerful and efficient tool for TD MS of proteins.
Cross-linking mass spectrometry (XL-MS) has become a very useful tool for studying protein complexes and interactions in living systems. It enables the investigation of many large and dynamic assemblies in their native state, providing an unbiased view of their protein interactions and restraints for integrative modeling. More researchers are turning toward trying XL-MS to probe their complexes of interest, especially in their native environments. However, due to the presence of other potentially higher abundant proteins, sufficient cross-links on a system of interest may not be reached to achieve satisfactory structural and interaction information. There are currently no rules for predicting whether XL-MS experiments are likely to work or not; in other words, if a protein complex of interest will lead to useful XL-MS data. Here, we show that a simple iBAQ (intensity-based absolute quantification) analysis performed from trypsin digest data can provide a good understanding of whether proteins of interest are abundant enough to achieve successful cross-linking data. Comparing our findings to large-scale data on diverse systems from several other groups, we show that proteins of interest should be at least in the top 20% abundance range to expect more than one cross-link found per protein. We foresee that this guideline is a good starting point for researchers who would like to use XL-MS to study their protein of interest and help ensure a successful cross-linking experiment from the beginning. Data are available via ProteomeXchange with identifier PXD045792.
In the last decade, MALDI-TOF Mass Spectrometry (MALDI-TOF MS) has been introduced and broadly accepted by clinical laboratory laboratories throughout the world as a powerful and efficient tool for rapid microbial identification. During the MALDI-TOF MS process, microbes are identified using either intact cells or cell extracts. The process is rapid, sensitive, and economical in terms of both labor and costs involved. Whilst MALDI-TOF MS is currently the gold-standard, it suffers from several shortcomings such as lack of direct information on antibiotic resistance, poor depth of analysis and insufficient discriminatory power for the distinction of closely related bacterial species or for reliably sub-differentiating isolates to the level of clones or strains. Thus, new approaches targeting proteins and allowing a better characterization of bacterial strains are strongly needed, if possible, on a very short time scale after sample collection in the hospital. Bottom-up proteomics (BUP) is a nice alternative to MALDI-TOF MS, offering the possibility for in-depth proteome analysis. Top-down proteomics (TDP) provides the highest molecular precision in proteomics, allowing the characterization of proteins at the proteoform level. A number of studies have already demonstrated the potential of these techniques in clinical microbiology. In this review, we will discuss the current state-of-the-art of MALDI-TOF MS for the rapid microbial identification and detection of resistance to antibiotics and describe emerging approaches, including bottom-up and top-down proteomics as well as ambient MS technologies.
Type IV pili (T4P) are prevalent, polymeric surface structures in pathogenic bacteria, making them ideal targets for effective vaccines. However, bacteria have evolved efficient strategies to evade type IV pili-directed antibody responses. Neisseria meningitidis are prototypical type IV pili-expressing Gram-negative bacteria responsible for life threatening sepsis and meningitis. This species has evolved several genetic strategies to modify the surface of its type IV pili, changing pilin subunit amino acid sequence, nature of glycosylation and phosphoforms, but how these modifications affect antibody binding at the structural level is still unknown. Here, to explore this question, we determine cryo-electron microscopy (cryo-EM) structures of pili of different sequence types with sufficiently high resolution to visualize posttranslational modifications. We then generate nanobodies directed against type IV pili which alter pilus function in vitro and in vivo. Cyro-EM in combination with molecular dynamics simulation of the nanobody-pilus complexes reveals how the different types of pili surface modifications alter nanobody binding. Our findings shed light on the impressive complementarity between the different strategies used by bacteria to avoid antibody binding. Importantly, we also show that structural information can be used to make informed modifications in nanobodies as countermeasures to these immune evasion mechanisms.
Proteoforms, which arise from post-translational modifications, genetic polymorphisms and RNA splice variants, play a pivotal role as drivers in biology. Understanding proteoforms is essential to unravel the intricacies of biological systems and bridge the gap between genotypes and phenotypes. By analysing whole proteins without digestion, top-down proteomics (TDP) provides a holistic view of the proteome and can decipher protein function, uncover disease mechanisms and advance precision medicine. This Primer explores TDP, including the underlying principles, recent advances and an outlook on the future. The experimental section discusses instrumentation, sample preparation, intact protein separation, tandem mass spectrometry techniques and data collection. The results section looks at how to decipher raw data, visualize intact protein spectra and unravel data analysis. Additionally, proteoform identification, characterization and quantification are summarized, alongside approaches for statistical analysis. Various applications are described, including the human proteoform project and biomedical, biopharmaceutical and clinical sciences. These are complemented by discussions on measurement reproducibility, limitations and a forward-looking perspective that outlines areas where the field can advance, including potential future applications. Proteoforms can be investigated using top-down proteomics, a technique that analyses whole proteins without previous digestion. This Primer introduces top-down proteomics, exploring mass spectrometry experimental methods, sample preparation, data analysis and applications in understanding human disease.
Periostin is a matricellular protein encoded by the POSTN gene that is alternatively spliced to produce ten different periostin isoforms with molecular weights ranging from 78 to 91 kDa. It is known to promote fibrillogenesis, organize the extracellular matrix, and bind integrin-receptors to induce cell signaling. As well as being a key component of the wound healing process, it is also known to participate in the pathogenesis of different diseases including atopic dermatitis, asthma, and cancer. In both health and disease, the functions of the different periostin isoforms are largely unknown. The ability to precisely determine the isoform profile of a given human sample is fundamental for characterizing their functional significance. Identification of periostin isoforms is most often carried out at the transcriptional level using RT-PCR based approaches, but due to high sequence homogeneity, identification on the protein level has always been challenging. Top-down proteomics, where whole proteins are measured by mass spectrometry, offers a fast and reliable method for isoform identification. Here we present a fully developed top-down mass spectrometry assay for the characterization of periostin splice isoforms at the protein level.
Traditionally, mass spectrometry (MS) output is the ion abundance plotted versus the ionic mass-to-charge ratio m/z. While employing only commercially available equipment, Charge Determination Analysis (CHARDA) adds a third dimension to MS, estimating for individual peaks their charge states z starting from z = 1 and color coding z in m/z spectra. CHARDA combines the analysis of ion signal decay rates in the time-domain data (transients) in Fourier transform (FT) MS with the interrogation of mass defects (fractional mass) of biopolymers. Being applied to individual isotopic peaks in a complex protein tandem (MS/MS) data set, CHARDA aids peptide mass spectra interpretation by facilitating charge-state deconvolution of large ionic species in crowded regions, estimating z even in the absence of an isotopic distribution (e.g., for monoisotopic mass spectra). CHARDA is fast, robust, and consistent with conventional FTMS and FTMS/MS data acquisition procedures. An effective charge-state resolution Rz ≥ 6 is obtained with the potential for further improvements.
Monoclonal antibodies (mAbs) have established themselves as the leading biopharmaceutical therapeutic modality. Once the developability of a mAb drug candidate has been assessed, an important step is to check its in vivo stability through pharmacokinetics (PK) studies. The gold standard is ligand-binding assay (LBA) and liquid chromatography-mass spectrometry (LC-MS) performed at the peptide level (bottom-up approach). However, these analytical techniques do not allow to address the different mAb proteoforms that can arise from biotransformation. In recent years, top-down and middle-down mass spectrometry approaches have gained popularity to characterize proteins at the proteoform level but are not yet widely used for PK studies. We propose here a workflow based on an automated immunocapture followed by top-down and middle-down liquid chromatography-tandem mass spectrometry (LC-MS/MS) approaches to characterize mAb proteoforms spiked in mouse plasma. We demonstrate the applicability of our workflow on a large concentration range using pembrolizumab as a model. We also compare the performance of two state-of-the-art Orbitrap platforms (Tribrid Eclipse and Exploris 480) for these studies. The added value of our workflow for an accurate and sensitive characterization of mAb proteoforms in mouse plasma is highlighted.
Motivation: There are several well-established paradigms for identifying and pinpointing discriminative peptides/ proteins using shotgun proteomic data; examples are peptide-spectrum matching, de novo sequencing, open searches, and even hybrid approaches. Such an arsenal of complementary paradigms can provide deep data coverage, albeit some unidentified discriminative peptides remain. Results: We present DiagnoMass, software tool that groups similar spectra into spectral clusters and then shortlists those clusters that are discriminative for biological conditions. DiagnoMass then communicates with proteomic tools to attempt the identification of such clusters. We demonstrate the effectiveness of DiagnoMass by analyzing proteomic data from Escherichia coli, Salmonella, and Shigella, listing many high-quality discriminative spectral clusters that had thus far remained unidentified by widely adopted proteomic tools. DiagnoMass can also classify proteomic profiles. We anticipate the use of DiagnoMass as a vital tool for pinpointing biomarkers. Availability: DiagnoMass and related documentation, including a usage protocol, are available at http://www. diagnomass.com.
Upon activation, vinculin reinforces cytoskeletal anchorage during cell adhesion. Activating ligands classically disrupt intramolecular interactions between the vinculin head and tail domains that bind to actin filaments. Here, we show that Shigella IpaA triggers major allosteric changes in the head domain, leading to vinculin homo-oligomerization. Through the cooperative binding of its three vinculin-binding sites (VBSs), IpaA induces a striking reorientation of the D1 and D2 head subdomains associated with vinculin oligomerization. IpaA thus acts as a catalyst producing vinculin clusters that bundle actin at a distance from the activation site and trigger the formation of highly stable adhesions resisting the action of actin relaxing drugs. Unlike canonical activation, vinculin homo-oligomers induced by IpaA appear to keep a persistent imprint of the activated state in addition to their bundling activity, accounting for stable cell adhesion independent of force transduction and relevant to bacterial invasion.
Introduction: Bordetella pertussis still circulates worldwide despite vaccination. Fimbriae are components of some acellular pertussis vaccines. Population fluctuations of B. pertussis fimbrial serotypes (FIM2 and FIM3) are observed, and fim3 alleles (fim3-1 [clade 1] and fim3-2 [clade 2]) mark a major phylogenetic subdivision of B. pertussis. Objectives: To compare microbiological characteristics and expressed protein profiles between fimbrial serotypes FIM2 and FIM3 and genomic clades. Methods: A total of 19 isolates were selected. Absolute protein abundance of the main virulence factors, autoagglutination and biofilm formation, bacterial survival in whole blood, induced blood cell cytokine secretion, and global proteome profiles were assessed. Results: Compared to FIM3, FIM2 isolates produced more fimbriae, less cellular pertussis toxin subunit 1 and more biofilm, but auto-agglutinated less. FIM2 isolates had a lower survival rate in cord blood, but induced higher levels of IL-4, IL-8 and IL-1b secretion. Global proteome comparisons uncovered 15 differentially produced proteins between FIM2 and FIM3 isolates, involved in adhesion and metabolism of metals. FIM3 isolates of clade 2 produced more FIM3 and more biofilm compared to clade 1. Conclusion: FIM serotype and fim3 clades are associated with proteomic and other biological differences, which may have implications on pathogenesis and epidemiological emergence. (c) 2023 Institut Pasteur. Published by Elsevier Masson SAS. All rights reserved.
In antibody-based drug research, a complete characterization of antibody proteoforms covering both the amino acid sequence and all posttranslational modifications remains a major concern. The usual mass spectrometry-based approach to achieve this goal is bottom-up proteomics, which relies on the digestion of antibodies but does not allow the diversity of proteoforms to be assessed. Middle-down and top-down approaches have recently emerged as attractive alternatives but are not yet mastered and thus used in routine by many analytical chemistry laboratories. The work described here aims at providing guidelines to achieve the best sequence coverage for the fragmentation of intact light and heavy chains generated from a simple reduction of intact antibodies using Orbitrap mass spectrometry. Three parameters were found crucial to this aim: the use of an electron-based activation technique, the multiplex selection of precursor ions of different charge states, and the combination of replicates.
Generating top-down tandem mass spectra (MS/MS) from complex mixtures of proteoforms benefits from improvements in fractionation, separation, fragmentation, and mass analysis. The algorithms to match MS/MS to sequences have undergone a parallel evolution, with both spectral alignment and match-counting approaches producing high-quality proteoform-spectrum matches (PrSMs). This study assesses state-of-the-art algorithms for top-down identification (ProSight PD, TopPIC, MSPathFinderT, and pTop) in their yield of PrSMs while controlling false discovery rate. We evaluated deconvolution engines (ThermoFisher Xtract, Bruker AutoMSn, Matrix Science Mascot Distiller, TopFD, and FLASHDeconv) in both ThermoFisher Orbitrap-class and Bruker maXis Q-TOF data (PXD033208) to produce consistent precursor charges and mass determinations. Finally, we sought post-translational modifications (PTMs) in proteoforms from bovine milk (PXD031744) and human ovarian tissue. Contemporary identification workflows produce excellent PrSM yields, although approximately half of all identified proteoforms from these four pipelines were specific to only one workflow. Deconvolution algorithms disagree on precursor masses and charges, contributing to identification variability. Detection of PTMs is inconsistent among algorithms. In bovine milk, 18% of PrSMs produced by pTop and TopMG were singly phosphorylated, but this percentage fell to 1% for one algorithm. Applying multiple search engines produces more comprehensive assessments of experiments. Top-down algorithms would benefit from greater interoperability.