The structure and function of many proteins are regulated post-translationally through glycan attachment. These glycans, assembled via competing enzymatic reactions, generate diverse glycoform populations - variants sharing a protein backbone but differing in glycan structures. While current analyses often focus on individual glycoforms, we demonstrate that population-level glycoform analysis - integrating spectral, biosynthetic, and physicochemical relationships - reveals new insights into glycoprotein regulation. Applied to immunoglobulin subclasses and antithrombin III (AT3), this approach provides comprehensive coverage of glycoform repertoires from human and murine plasma and biopharmaceuticals. It also enables sensitive quantification of glycosylation changes arising from in vitro manipulations or in vivo infections. Finally, we introduce a statistical framework adapted from ecological biodiversity studies, revealing that both IgG and AT3 exhibit skewed glycoform distributions shaped by biosynthetic constraints and degradation. Our findings demonstrate the added value of population-level glycoform analysis in understanding protein function and regulation through glycosylation.
AIM:Risk assessment tools for periodontitis lack sufficient precision to predict disease progression. The aim of this exploratory study was to evaluate proteolytic activity (PA) in gingival crevicular fluid in relation to clinically assessed progression risk, microbial community composition and host proteome. MATERIALS AND METHODS:Samples were collected from 109 patients with periodontitis and stratified into high and low PA groups using a fluorescent substrate. Microbial composition was analysed via 16S rRNA gene sequencing, and host proteomic profiles were determined using mass spectrometry. RESULTS:Forty-one of 42 patients in the high PA group were clinically classified as high-risk, while the low-activity group included both low- and high-risk individuals. High-activity samples displayed greater microbial diversity and enrichment of proteolytic species including Fusobacterium nucleatum, Porphyromonas gingivalis and Treponema denticola. Of the 235 differentially expressed host proteins, high-activity samples had elevated markers of inflammation and neutrophil-mediated immunity, whereas low-activity samples were enriched in protease inhibitors, indicating more effective regulation of tissue degradation. CONCLUSIONS:Small-volume gingival crevicular fluid sampling allows integrated analysis of microbial and host signatures that provide insight into the pathobiology of periodontitis. PA stratifies patients according to these microbial and host proteome signatures.
High-throughput proteomics has enabled detailed characterization of molecular states across health and disease. However, biological systems are inherently dynamic and methods for reconstructing continuous proteome changes remain limited. Here, we introduce proteome velocity, a framework for inferring continuous proteome trajectories from cross-sectional or sparsely sampled proteomics data using flow matching, in which a neural network learns velocity fields over proteome space. Proteome velocity estimates how rapidly and in which direction protein abundances change along a biological progression, such as disease. In mouse sepsis, covariate-conditioned velocity models resolved tissue- and pathogen-specific proteome trajectories and identified inflammatory proteins with distinct temporal activation patterns across infection routes and organ systems. In clinical COVID-19 plasma proteomes, inferred trajectories separated into distinct velocity programs associated with disease severity. These results show how generative trajectory models can transform cross-sectional proteomics data into interpretable, protein-resolved representations of molecular progression.
Chronic obstructive pulmonary disease (COPD) is a progressive condition characterized by airway remodeling, including emphysema and fibrosis. Proteoglycans and their glycosaminoglycan (GAG) chains are key components of the extracellular matrix and may be altered as the disease advances. This study analyzed lung tissue from COPD patients (GOLD stages II-III and IV), non-COPD smokers, and non-smokers to assess proteoglycan and GAG changes. While LC-MS revealed no alterations in chondroitin/dermatan sulfate (CS/DS) or heparan sulfate (HS) proteoglycan core proteins, the total GAG level increased in GOLD II-IV patients. HS displayed increased N- and 2-O-sulfation in GOLD IV, while CS/DS levels and 4-O-sulfation were enhanced across GOLD II-IV. These findings were supported by transcriptomic data indicating upregulation of CHST11, the main CS/DS 4-O-sulfotransferase. In line with previous findings, TGF-β signaling was shown to be enriched in COPD patients and to regulate the CHST11 expression. These results were confirmed by TGF-β stimulation of lung fibroblasts showing increased CS/DS levels, 4-O-sulfation, and CHST11 expression. In conclusion, COPD is associated with disease-stage-specific changes in GAG sulfation, particularly enhanced CS/DS 4-O-sulfation that is likely to be driven by TGF-β. These alterations may contribute to extracellular matrix remodeling and represent potential targets for therapeutic intervention to mitigate disease progression.
Bartonella schoenbuchensis is suspected to cause deer ked dermatitis and febrile diseases in humans. Deer keds (Lipoptena cervi), which infest cervids (e.g., roe deer, fallow deer), are discussed as potential vectors for B. schoenbuchensis. We analyzed the seroprevalence of anti-B. schoenbuchensis immunoglobulin G (IgG) antibodies in sera of forest workers (FW; n = 82) compared to control sera of non-forest workers (NFW; n = 118) from North Rhine-Westphalia, Germany. For this purpose, an immunofluorescence assay (IFA) using Vero E6 cells infected with B. schoenbuchensis was established, and serum titers were assessed. Whole cell lysate of B. schoenbuchensis was introduced for analysis of seroreactivity by western blotting. Immunodominant proteins were identified by liquid chromatography–tandem mass spectrometry. When using human sera, 54.9
Designing novel proteins with desired characteristics remains a significant challenge due to the large sequence space and the complexity of sequence-function relationships. Efficient exploration of this space to identify sequences that meet specific design criteria is crucial for advancing therapeutics and biotechnology. Here, we present BoGA (Bayesian Optimization Genetic Algorithm), a framework that combines evolutionary search with Bayesian optimization to efficiently navigate the sequence space. By integrating a genetic algorithm as a stochastic proposal generator within a surrogate modeling loop, BoGA prioritizes candidates based on prior evaluations and surrogate model predictions, enabling data-efficient optimization. We demonstrate the utility of BoGA through benchmarking on sequence and structure design tasks, followed by its application in designing peptide binders against pneumolysin, a key virulence factor of \textit{Streptococcus pneumoniae}. BoGA accelerates the discovery of high-confidence binders, demonstrating the potential for efficient protein design across diverse objectives. The algorithm is implemented within the BoPep suite and is available under an MIT license at \href{https://github.com/ErikHartman/bopep}{GitHub}.
Abstract Streptococcus pneumoniae remains a major public health concern, largely due to the limited serotype coverage and other constraints of pneumococcal conjugate vaccines, as well as the increasing antimicrobial resistance among circulating strains. In the pursuit of protein-based pneumococcal vaccines, pneumolysin (PLY), a secreted multifunctional cholesterol-dependent cytolysin, represents a promising target. To date, the relationship between the B-cell epitope landscape and neutralising PLY-specific antibody responses has remained elusive, hindering the rational design of effective PLY-based vaccines. Using a panel of PLY-specific monoclonal antibodies, functional assays, multimodal protein mass spectrometry and data-driven computational modelling, we mapped the structural epitope landscape of native PLY and linked epitope-paratope interactions to neutralising potency. We further refined and structurally characterised a protective cross-species epitope conserved among homologous cholesterol-dependent cytolysins (CDC). This epitope provides a promising foundation for rational, epitope-focused vaccine design, offering a pathway toward species-independent vaccines targeting the CDC protein superfamily.
Protein-protein interactions (PPIs) are central to cellular processes and host-pathogen dynamics across all domains of life, yet comprehensive interactome mapping remains challenging at the proteome scale. Experimental approaches provide only partial coverage, while existing computational methods often lack generalizability across species or are too resource-intensive for large-scale screening. Here, we introduce ppIRIS (protein-protein Interaction Regression via Iterative Siamese networks), a lightweight deep learning framework that integrates evolutionary and structural embeddings to predict PPIs directly from sequence. Evaluated on multi-species benchmarks, ppIRIS achieves state-of-the-art accuracy while enabling proteome-wide screening in minutes. Trained on curated bacterial datasets and applied to the Group A Streptococcus (GAS) proteome, ppIRIS identified functional clusters associated with virulence pathways, such as nutrient transport, stress response, and metal scavenging. Extending to cross-species prediction, ppIRIS recovered 56.2% of known GAS-human plasma interactions with enrichment in complement, coagulation, and protease inhibition pathways. Experimental validation confirmed novel predictions, demonstrating the applicability of ppIRIS for systematic discovery of bacterial and cross-species PPIs. The model together with a Google Colaboratory is freely available at github.com/lupiochi/ppIRIS.
Protein degradation is a regulated process that reshapes the proteome and generates bioactive peptides. Peptidomics and degradomics enables large-scale measurement of these peptides, yet most data analyses approaches treat peptides as isolated endpoints rather than intermediates produced by sequential cleavage. Here, we introduce degradation graphs, a probabilistic framework that represents proteolysis as a directed acyclic network of cleavage events with explicit absorption. From single-snapshot peptidomes, we infer graph weights by gradient descent or linear-flow optimization, quantify flows through branches and bottlenecks, and correct a core bias in conventional quantification. Across three biological datasets, failure to model downstream trimming leads to 3-4-fold underestimation of upstream proteolytic activity. Moreover, degradation graphs provide graph-structured features that enable machine learning models to capture protease-specific signatures from both graph topology and sequence context. Taken together, these findings establish explicit degradation modeling as a practical approach to mechanistic and interpretable peptidomics, bridging the fields of degradomics and peptidomics.
The endogenous peptidome is the complete set of naturally occurring peptides, some of which are involved in processes like intercellular communication and in innate immunity. The peptidome is also a potential source of uncharacterized peptide-based therapeutics, which has remained largely unexplored due to challenges in identifying bioactive peptides in vast peptidomic datasets. Here, we present a generalized computational pipeline that mines datasets for peptides with high affinity binding capacities. The approach employs protein-peptide docking, binding interface scoring and a deep ensemble model to dynamically learn the mapping between peptide embeddings and binding capacity. We demonstrate the utility of the pipeline by identifying endogenous peptide binders for the LPS-binding site of CD14 using experimentally defined wound fluid peptidomes. Promising candidates underwent conformational refinement and validation through all-atom molecular dynamics simulations. This pipeline offers a systematic way to uncover novel peptide binders within large peptidomes, paving the way for accelerated discovery of peptide-based therapeutics. ### Competing Interest Statement A.S. is a founder of in2cure AB, a parent company of Xinnate and Transient Pharma AB, companies that are developing therapies based on thrombin-derived peptides and variants. The TCP family, including stapled variants, is patent protected.
Background:A pilot study investigating proteomic profiles from 78 patients from the Target Temperature Management after Out-of-hospital Cardiac arrest (TTM) trial revealed 35 proteins associated to functional outcome, and six proteins associated to targeted temperature management at 33 °C. We present the protocol for a study investigating proteomic profiles in the full cohort of the TTM-trial biobank. The aim is to stratify protein profiles based on survival, functional outcome, targeted temperature management, and MIRACLE2 score in order to search for potential novel biomarkers. Methods:All patients with available serum samples at 24, 48, and/or 72 h after return of spontaneous circulation (N = 682 patients and N = 1882 samples) will be included in the liquid chromatography and tandem mass spectrometry analysis using diaPASEF, combining data-independent-acquisition of spectra with parallel accumulation-serial fragmentation. Statistical analysis will include data normalisation, exploratory principal component analysis, and differential expression analysis. Changes in serum protein abundance will be analysed according to survival and binary functional outcome (modified Rankin Scale 0-3 vs. 4-6) at six-months after randomisation, randomisation to target temperature of 33 °C or 36 °C, and the MIRACLE2 score. Secondary stratifications will include sex, age, time to return of spontaneous circulation, shockable vs. non-shockable initial rhythm, circulatory shock on admission, and presumed cause of death. Conclusion:This prospective study will provide information about proteomic profiles after cardiac arrest and may give insight for identification of novel biomarkers for prediction of outcome.
Streptococcus pyogenes (Group A Streptococcus , GAS) is a significant human pathogen for which no licensed vaccine is currently available. Here, we report a de novo designed epitope-centric protein-based nanoparticle vaccine against GAS. By integrating structural mass spectrometry techniques and deep learning approaches, we re-engineered a protective epitope (D3m) present in domain 3 of streptolysin O, a prominent pore-forming toxin produced by GAS. D3m was displayed on the surface of a self-assembling icosahedral nanoparticle (D3m-NP) to enhance epitope presentation and immunogenicity. Mice immunised with D3m-NP mounted haemolysis-neutralising titres and displayed a more uniform, epitope-centric antibody response than those receiving the community-standard detoxified full-length streptolysin O. Our findings highlight a promising strategy for GAS vaccine development by combining multimodal protein mass spectrometry, protein design and a versatile protein-based nanoparticle vaccine platform. ### Competing Interest Statement D.T., E.H., L.M., and J.M. disclose a pending patent application associated with the research findings reported in this article. Sten K Johnsons Stiftelse, 802477-7123 Knut and Alice Wallenberg Foundation, https://ror.org/004hzzk67, KAW 2017.0271, KAW 2016.0023, KAW 2019.0353, KAW 2020.0299 Swedish Research Council, https://ror.org/03zttf063, 2019-01646, 2018-05795 Alfred Österlunds Stiftelse
Persistent bacterial airway infection is a hallmark feature of cystic fibrosis (CF). Achromobacter spp. are gram-negative rods that can cause persistent airway infection in people with CF (pwCF), but the knowledge of host immune responses to these bacteria is limited. The aim of this study was to investigate if patients develop antibodies against Achromobacter xylosoxidans, the most common Achromobacter species, and to identify the bacterial antigens that induce specific IgG responses. Seven serum samples from pwCF with Achromobacter infection were screened for antibodies against bacteria in an ELISA coated with A. xylosoxidans, A. insuavis, or Pseudomonas aeruginosa. Sera from pwCF with or without P. aeruginosa infection (n = 22 and 20, respectively) and healthy donors (n = 4) were included for comparison. Serum with high titers to A. xylosoxidans was selected for affinity purification of bacterial antigens using serum IgGs bound to protein G beads. The resulting IgG-antigen complexes were then analyzed using liquid chromatography-tandem mass spectrometry (LC-MS/MS). Selected antigens of interest were produced in recombinant form and used in an ELISA to confirm the results. Four of the seven patients with Achromobacter infection had serum antibodies against Achromobacter. Using patient serum-IgG for affinity purification of A. xylosoxidans proteins, we identified eight antigens. Three of these, which were not targeted by anti-P. aeruginosa antibodies, were expressed recombinantly for further validation: dihydrolipoyl dehydrogenase (DLD), type I secretion C-terminal target domain-containing protein, and domain of uncharacterized function 336 (DUF336). While specific IgG against all three recombinant antigens was confirmed in the patient serum with high titers against Achromobacter, DLD and DUF336 showed the least binding to serum IgG from pwCF without Achromobacter spp. infection. Using serum IgG affinity purification in combination with LC-MS/MS and confirming the results using ELISA against recombinant proteins, we have identified bacterial antigens from A. xylosoxidans.IMPORTANCEAchromobacter species are opportunistic pathogens that can cause airway infections in people with cystic fibrosis. In this patient population, persistent Achromobacter infection is associated with low lung function, but the knowledge about bacterial interactions with the host is currently limited. In this study, we identify protein antigens that induce specific antibody responses in the host. The identified antigens may potentially be useful in serological assays, serving as a complement to culturing methods for the diagnosis and surveillance of Achromobacter infection.
The Warburg effect, which describes the fermentation of glucose to lactate even in the presence of oxygen, is ubiquitous in proliferative mammalian cells, including cancer cells, but poses challenges for biopharmaceutical production as lactate accumulation inhibits cell growth and protein production. Previous efforts to eliminate lactate production in cells for bioprocessing have failed as lactate dehydrogenase is essential for cell growth. Here, we effectively eliminate lactate production in Chinese hamster ovary and in the human embryonic kidney cell line HEK293 by simultaneous knockout of lactate dehydrogenases and pyruvate dehydrogenase kinases, thereby removing a negative feedback loop that typically inhibits pyruvate conversion to acetyl-CoA. These cells, which we refer to as Warburg-null cells, maintain wild-type growth rates while producing negligible lactate, show a compensatory increase in oxygen consumption, near total reliance on oxidative metabolism, and higher cell densities in fed-batch cell culture. Warburg-null cells remain amenable for production of diverse biotherapeutic proteins, reaching industrially relevant titres and maintaining product glycosylation. The ability to eliminate lactate production may be useful for biotherapeutic production and provides a tool for investigating a common metabolic phenomenon. Hefzi et al. engineer cells to nearly eliminate lactate production and increase mitochondrial use of pyruvate by deleting both lactate dehydrogenases and pyruvate dehydrogenase kinases with the goal of circumventing pH and osmolarity issues that arise in bioreactor-based production of biopharmaceuticals.
This study showcases an integrative mass spectrometry-based strategy combining systems antigenomics and systems serology to characterize human antibodies in clinical samples. This strategy involves using antibodies circulating in plasma to affinity-enrich antigenic proteins in biochemically fractionated pools of bacterial proteins, followed by their identification and quantification using mass spectrometry. A selected subset of the identified antigens is then expressed recombinantly to isolate antigen-specific IgG, followed by characterization of the structural and functional properties of these antibodies. We focused on Group A streptococcus (GAS), a major human pathogen lacking an approved vaccine. The data shows that both healthy and GAS-infected individuals have circulating IgG against conserved streptococcal proteins, including toxins and virulence factors. The antigenic breadth of these antibodies remains relatively constant across healthy individuals but changes considerably in GAS bacteremia. Moreover, antigen-specific IgG analysis reveals individual variation in titers, subclass distributions, and Fc-signaling capacity, despite similar epitope and Fc-glycosylation patterns. Finally, we show that GAS antibodies may cross-react with Streptococcus dysgalactiae (SD), a bacterial pathogen that occupies similar niches and causes comparable infections. Collectively, our results highlight the complexity of GAS-specific antibody responses and the versatility of our methodology to characterize immune responses to bacterial pathogens.
Proteins underpin most biological function, and the ability to design them with tailored structures and properties is central to advances in biotechnology. Diffusion-based generative models have emerged as powerful tools for protein design, but steering them toward proteins with specified properties remains challenging. The Feynman-Kac (FK) framework provides a principled way to guide diffusion models using user-defined rewards. In this paper, we enable FK-based steering of RFdiffusion through the development of guiding potentials that leverage ProteinMPNN and structural relaxation to guide the diffusion process towards desired properties. We show that steering can be used to consistently improve predicted interface energetics and increase binder designability by 89.5%. Together, these results establish that diffusion-based protein design can be effectively steered toward arbitrary, non-differentiable objectives, providing a model-independent framework for controllable protein generation.
Protein–protein interactions (PPIs) are central to cellular processes and host–pathogen dynamics, yet bacterial interactomes remain poorly mapped, especially for extracellular effectors and cross-species interactions. Experimental approaches provide only partial coverage, while existing computational methods often lack generalizability or are too resource-intensive for proteome-scale application. Here, we introduce ppIRIS (protein–protein Interaction Regression via Iterative Siamese networks), a lightweight deep learning model that integrates evolutionary and structural embeddings to predict PPIs directly from sequence. Trained on curated bacterial datasets, ppIRIS achieves state-of-the-art accuracy across benchmarks while enabling proteome-wide screening in minutes. Applied to Group A Streptococcus (GAS), ppIRIS revealed functional clusters linked to virulence pathways, including nutrient transport, stress response, and metal scavenging. For host–pathogen predictions, ppIRIS recovered 56.2% of known GAS–human plasma interactions with enrichment in complement, coagulation, and protease inhibition pathways. Experimental validation confirmed novel predictions, demonstrating the applicability of ppIRIS for systematic discovery of bacterial and cross-species PPIs. The software is freely available at [github.com/lupiochi/ppIRIS][1]. ### Competing Interest Statement The authors have declared no competing interest. Agence Nationale de la Recherche, https://ror.org/00rbzpz17, ANR-22-CPJ2-0075-01, ANR-24-CE45-4243-01 [1]: http://github.com/lupiochi/ppIRIS
Recently, mass spectrometry based peptidomics studies have proven useful in the identification of biomarkers and bioactive peptide-based therapeutics. Here, we present a dataset comprised of temporal wound fluid peptidomics data from highly defined porcine models. Wound fluids from porcine wounds infected with Staphylococcus aureus and Pseudomonas aeruginosa, and uninfected controls, were sampled at different timepoints of the infection. Peptides were extracted from the samples, followed by liquid chromatography tandem mass spectrometry analysis in data dependent acquisition mode. The resulting spectra and searched files have been deposited to online repositories and made easily accessible to enable further investigations of the infected and uninfected wound fluid peptidome.
Sepsis is a life-threatening condition caused by a dysregulated host response to an infection and is a leading cause of death worldwide. The condition is variable, which in combination with insufficient clinical markers, makes it challenging to predict when infection will progress to sepsis and to categorize patients into homogeneous patient subgroups. In this study, we demonstrate the use of acoustic trapping to rapidly enrich extracellular vesicles (EVs) from minute volumes of blood plasma from experimental mouse models of sepsis infected with the Gram-positive pathogen Streptococcus pyogenes or the Gram-negative pathogen Escherichia coli. Using quantitative mass spectrometry-based proteomics, we characterized the proteome of EVs and plasma to demonstrate that the EVs expand the observable proteome in plasma, with an emphasis on cellular processes and signaling. In our models, systemic bacterial infection altered the EV and plasma proteomes differently, with a predominant effect on proteins related to leukocyte migration in the EVs and on metabolism in the plasma. Finally, we show that E. coli infection significantly impacted metabolism, whereas S. pyogenes infection mainly affected the inflammatory response and neutrophil degranulation in our models. Collectively, our findings demonstrate that the acoustic trap facilitates access to plasma EVs, which in turn provides additional biological information which was not obtained from the plasma proteome alone.