Motivation Mass spectrometry-based proteomics allows studying all proteins of a sample on a molecular level. However, mass spectra are noisy and contain complex patterns, making them inherently challenging to analyze with algorithmic approaches. In terms of the protein sequence landscape, most recent bottom-up MS-based proteomics studies consider either a diverse pool of post-translational modifications, employ large databases-as in metaproteomics or proteogenomics, study multiple isoforms of proteins, include unspecific cleavage sites or even combinations thereof. All this makes peptide and protein identifications challenging.Results Here, we present a foundation model, called yHydra, that jointly embeds spectra and peptides. This allows us to implement various downstream tasks and search modes in Euclidean space. We implement an open search which allows querying multiple ten-thousands of spectra against millions of peptides. Furthermore, we implement an error-tolerant search for identifying additional proteoforms that are not included in off-the-shelf reference proteomes. Our foundation model provides meaningful embeddings, as we interpret learned peptide embeddings in comparison to the peptide's physico-chemical properties. Hydra's open search, assigns delta masses to each identification which allows to unrestrictedly characterize post-translational modifications. The error-tolerant mode of yHydra can be used as post-processing to existing search engines or as a standalone. yHydra is evaluated on several real life data sets for the identification of modified peptide sequences and shows up to 25% increase in peptide identification at constant false discovery rate compared to the current state-of-the-art.Availability and Implementation Code is available on Gitlab: https://gitlab.com/dacs-hpi/yHydra, and https://gitlab.com/dacs-hpi/yHydra_train.
BACKGROUNDAccurate prognostic assays for COVID-19 represent an unmet clinical need. We sought to identify and validate early parsimonious transcriptomic signatures that accurately predict fatal outcomes.METHODSWe studied 894 patients enrolled in the prospective, multicenter Immunophenotyping Assessment in a COVID-19 Cohort (IMPACC) with peripheral blood mononuclear cells (PBMC) and nasal swabs collected within 48 hours of admission. Host gene expression was measured with RNA-Seq. We trained parsimonious prognostic classifiers incorporating host gene expression, age, and SARS-CoV-2 viral load to predict 28-day mortality in 70% of the cohort. Classifier performance was determined in the remaining 30% and externally validated in a contemporary COVID-19 cohort (n = 137) with vaccinated patients.RESULTSFatal COVID-19 was characterized by 4,189 differentially expressed genes in the peripheral blood. A COVID-specific 3-gene peripheral blood classifier (CD83, ATP1B2, DAAM2) combined with age and SARS-CoV-2 viral load achieved an area under the receiver operating characteristic curve (AUC) of 0.88 (95% CI, 0.82-0.94). A 3-gene nasal classifier (SLC5A5, CD200R1, FCER1A), in comparison, yielded an AUC of 0.74 (95% CI, 0.64-0.83). Notably, OLAH, the most strongly upregulated gene in both PBMC and nasal swab and recently implicated in severe viral infection pathogenesis, yielded AUCs of 0.86 (0.79-0.93) and 0.78 (95% CI, 0.69-0.86), respectively. Both peripheral blood classifiers demonstrated comparable performance in an independent contemporary cohort of vaccinated patients (AUCs 0.74-0.80).CONCLUSIONOur parsimonious blood- and nasal-based classifiers accurately predicted COVID-19 mortality and merit further study as accessible prognostic tools to guide triage, resource allocation, and early therapeutic interventions.FUNDINGNIH: 5R01AI135803-03, R35HL140026, 5U19AI118608-04, 5U19AI128910-04, 4U19AI090023-11, 4U19AI118610-06, R01AI145835-01A1S1, 5U19AI062629-17, 5U19AI057229-17, 5U19AI125357-05, 5U19AI128913-03, 3U19AI077439-13, 5U54AI142766-03, 5R01AI104870-07, 3U19AI089992-09, 3U19AI128913-03, 5T32DA018926-18, and K0826161611. National Institute of Allergy and Infectious Diseases, NIH: 3U19AI1289130, U19AI128913-04S1, and R01AI122220. National Center for Advancing Translational Sciences, NIH: UM1TR004528. The National Science Foundation: DMS2310836. The Chan Zuckerberg Biohub San Francisco.
Predicting mortality risk in patients with COVID-19 remains challenging, and accurate prognostic assays represent a persistent unmet clinical need. We aimed to identify and validate parsimonious transcriptomic signatures that accurately predict fatal outcomes within 48 hours of hospitalization. We studied 894 patients hospitalized for COVID-19 across 20 US hospitals and enrolled in the prospective Immunophenotyping Assessment in a COVID-19 Cohort (IMPACC) with peripheral blood mononuclear cells (PBMC) and nasal swabs collected within 48 hours of admission. Host gene expression was assessed by RNA sequencing, nasal SARS-CoV-2 viral load was measured by RT-qPCR, and mortality was assessed at 28 days. We first defined transcriptional signatures and biological features of fatal COVID-19, which we compared against mortality signatures from an independent cohort of patients with non-COVID-19 sepsis (n=122). Using least absolute shrinkage and selection operator (LASSO) regression in 70% of the COVID-19 cohort, we trained parsimonious prognostic classifiers incorporating host gene expression, age, and viral load. The performance of single and three-gene classifiers was then determined in the remaining 30% of the cohort and subsequently externally validated in an independent, contemporary COVID-19 cohort (n=137) with vaccinated patients. Fatal COVID-19 was characterized by 4189 differentially expressed genes in the peripheral blood, representing marked upregulation of neutrophil degranulation, erythrocyte gas exchange, and heme biosynthesis pathways, juxtaposed against downregulation of adaptive immune pathways. Only 7.6% of mortality-associated genes overlapped between COVID-19 and sepsis due to other causes. A COVID-specific three-gene peripheral blood classifier ( CD83, ATP1B2, DAAM2 ) combined with age and SARS-CoV-2 viral load achieved an area under the receiver operating characteristic curve (AUC) of 0.88 (95% CI 0.82–0.94). A three-gene nasal classifier ( SLC5A5, CD200R1 , FCER1A ), in comparison, yielded an AUC of 0.74 (95% CI 0.64-0.83). Notably the expression of OLAH alone, a gene recently implicated in severe viral infection pathogenesis, yielded an AUC of 0.86 (0.79–0.93). Both peripheral blood classifiers demonstrated comparable performance in vaccinated patients from an independent external validation cohort (AUCs 0.74– 0.80). A three-gene peripheral blood signature, as well as OLAH alone, accurately predict COVID-19 mortality early in hospitalization, including in vaccinated patients. These parsimonious blood- and nasal-based classifiers merit further study as accessible prognostic tools to guide triage, resource allocation, and early therapeutic interventions in COVID-19.
Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) infection is characterized by highly heterogeneous manifestations ranging from asymptomatic cases to death for still incompletely understood reasons. As part of the IMmunoPhenotyping Assessment in a COVID-19 Cohort study, we mapped the plasma proteomes of 1117 hospitalized patients with COVID-19 from 15 hospitals across the United States. Up to six samples were collected within ~28 days of hospitalization resulting in one of the largest COVID-19 plasma proteomics cohorts with 2934 samples. Using perchloric acid to deplete the most abundant plasma proteins allowed for detecting 2910 proteins. Our findings show that increased levels of neutrophil extracellular trap and heart damage markers are associated with fatal outcomes. Our analysis also identified prognostic biomarkers for worsening severity and death. Our comprehensive longitudinal plasma proteomics study, involving 1117 participants and 2934 samples, allowed for testing the generalizability of the findings of many previous COVID-19 plasma proteomics studies using much smaller cohorts.
Age is a major risk factor for severe coronavirus disease 2019 (COVID-19), yet the mechanisms behind this relationship have remained incompletely understood. To address this, we evaluated the impact of aging on host immune response in the blood and the upper airway, as well as the nasal microbiome in a prospective, multicenter cohort of 1031 vaccine-naïve patients hospitalized for COVID-19 between 18 and 96 years old. We performed mass cytometry, serum protein profiling, anti–severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) antibody assays, and blood and nasal transcriptomics. We found that older age correlated with increased SARS-CoV-2 viral abundance upon hospital admission, delayed viral clearance, and increased type I interferon gene expression in both the blood and upper airway. We also observed age-dependent up-regulation of innate immune signaling pathways and down-regulation of adaptive immune signaling pathways. Older adults had lower naïve T and B cell populations and higher monocyte populations. Over time, older adults demonstrated a sustained induction of pro-inflammatory genes and serum chemokines compared with younger individuals, suggesting an age-dependent impairment in inflammation resolution. Transcriptional and protein biomarkers of disease severity differed with age, with the oldest adults exhibiting greater expression of pro-inflammatory genes and proteins in severe disease. Together, our study finds that aging is associated with impaired viral clearance, dysregulated immune signaling, and persistent and potentially pathologic activation of pro-inflammatory genes and proteins.
We introduce a cost-effective, robust high-throughput–compatible plasma depletion method enabling in-depth profiling of plasma that detects >1300 proteins per run with a throughput of 60 samples per day. The method has been fully validated by processing >3000 samples with no apparent batch effect at a cost for the depletion step of ~$2.5 per sample.
BackgroundMouse models that overexpress human mutant Tau (P301S and P301L) are commonly used in preclinical studies of Alzheimer's Disease (AD) and while several drugs showed therapeutic effects in these mice, they were ineffective in humans. This leads to the question to which extent the murine models reflect human Tau pathology on the molecular level.MethodsWe isolated insoluble, aggregated Tau species from two common AD mouse models during different stages of disease and characterized the modification landscape of the aggregated Tau using targeted and untargeted mass spectrometry-based proteomics. The results were compared to human AD and to human patients that suffered from early onset dementia and that carry the P301L Tau mutation.ResultsBoth mouse models accumulate insoluble Tau species during disease. The Tau aggregation is driven by progressive phosphorylation within the proline rich domain and the C-terminus of the protein. This is reflective of early disease stages of human AD and of the pathology of dementia patients carrying the P301L Tau mutation. However, Tau ubiquitination and acetylation, which are important to late-stage human AD are not represented in the mouse models.ConclusionAD mouse models that overexpress human Tau using risk mutations are a suitable tool for testing drug candidates that aim to intervene in the early formation of insoluble Tau species promoted by increased phosphorylation of Tau.
OBJECTIVE:To determine if oral secretions (OS) can be used as a noninvasively collected body fluid, in lieu of tracheal aspirates (TA), to track respiratory status and predict bronchopulmonary dysplasia (BPD) development in infants born <32 weeks. STUDY DESIGN:This was a retrospective, single center cohort study that included data and convenience samples from week-of-life (WoL) 3 from 2 independent preterm infant cohorts. Using previously banked samples, we applied our sample-sparing, high-throughput proteomics technology to compare OS and TA proteomes in infants born <32 weeks admitted to the Neonatal Intensive Care Unit (NICU) (Cohort 1; n = 23 infants). In a separate similar cohort, we mapped the BPD-associated changes in the OS proteome (Cohort 2; n = 17 infants including 8 with BPD). RESULTS:In samples collected during the first month of life, we identified 607 proteins unique to OS, 327 proteins unique to TA, and 687 overlapping proteins belonging to pathways involved in immune effector processes, neutrophil degranulation, leukocyte mediated immunity, and metabolic processes. Furthermore, we identified 37 OS proteins that showed significantly differential abundance between BPD cases and controls: 13 were associated with metabolic and immune dysregulation, 10 of which (eg, SERPINC1, CSTA, BPI) have been linked to BPD or other prematurity-related lung disease based on blood or TA investigations, but not OS. CONCLUSIONS:OS are a noninvasive, easily accessible alternative to TA and amenable to high-throughput proteomic analysis in preterm newborns. OS samples hold promise to yield actionable biomarkers of BPD development, particularly for prospective categorization and timely tailored treatment of at-risk infants with novel therapies.
The IMPACC cohort, composed of >1,000 hospitalized COVID-19 participants, contains five illness trajectory groups (TGs) during acute infection (first 28 days), ranging from milder (TG1-3) to more severe disease course (TG4) and death (TG5). Here, we report deep immunophenotyping, profiling of >15,000 longitudinal blood and nasal samples from 540 participants of the IMPACC cohort, using 14 distinct assays. These unbiased analyses identify cellular and molecular signatures present within 72 h of hospital admission that distinguish moderate from severe and fatal COVID-19 disease. Importantly, cellular and molecular states also distinguish participants with more severe disease that recover or stabilize within 28 days from those that progress to fatal outcomes (TG4 vs. TG5). Furthermore, our longitudinal design reveals that these biologic states display distinct temporal patterns associated with clinical outcomes. Characterizing host immune responses in relation to heterogeneity in disease course may inform clinical prognosis and opportunities for intervention.
To develop therapies for Alzheimer's disease, we need accurate in vivo diagnostics. Multiple proteomic studies mapping biomarker candidates in cerebrospinal fluid (CSF) resulted in little overlap. To overcome this shortcoming, we apply the rarely used concept of proteomics meta-analysis to identify an effective biomarker panel. We combine ten independent datasets for biomarker identification: seven datasets from 150 patients/controls for discovery, one dataset with 20 patients/controls for down-selection, and two data -sets with 494 patients/controls for validation. The discovery results in 21 biomarker candidates and down -selection in three, to be validated in the two additional large-scale proteomics datasets with 228 diseased and 266 control samples. This resulting 3-protein biomarker panel differentiates Alzheimer's disease (AD) from controls in the two validation cohorts with areas under the receiver operating characteristic curve (AUROCs) of 0.83 and 0.87, respectively. This study highlights the value of systematically re-analyzing pre-viously published proteomics data and the need for more stringent data deposition.
Herpesviruses have complex mechanisms enabling infection of the human CNS and evasion of the immune system, allowing for indefinite latency in the host. Herpesvirus infections can cause severe complications of the central nervous system (CNS). Here, we provide a novel characterization of cerebrospinal fluid (CSF) proteomes from patients with meningitis or encephalitis caused by human herpes simplex virus 1 (HSV-1), which is the most prevalent human herpesvirus associated with the most severe morbidity. The CSF proteome was compared with those from patients with meningitis or encephalitis due to human herpes simplex virus 2 (HSV-2) or varicella-zoster virus (VZV, also known as human herpesvirus 3) infections. Virus-specific differences in CSF proteomes, most notably elevated 14-3-3 family proteins and calprotectin (i.e., S100-A8 and S100-A9), were observed in HSV-1 compared to HSV-2 and VZV samples, while metabolic pathways related to cellular and small molecule metabolism were downregulated in HSV-1 infection. Our analyses show the feasibility of developing CNS proteomic signatures of the host response in alpha herpes infections, which is paramount for targeted studies investigating the pathophysiology driving virus-associated neurological disorders, developing biomarkers of morbidity, and generating personalized therapeutic strategies.
Combining robust proteomics instrumentation with high-throughput enabling liquid chromatography (LC) systems (e.g., timsTOF Pro and the Evosep One system, respectively) enabled mapping the proteomes of 1000s of samples. Fragpipe is one of the few computational protein identification and quantification frameworks that allows for the time-efficient analysis of such large data sets. However, it requires large amounts of computational power and data storage space that leave even state-of-the-art workstations underpowered when it comes to the analysis of proteomics data sets with 1000s of LC mass spectrometry runs. To address this issue, we developed and optimized a Fragpipe-based analysis strategy for a high-performance computing environment and analyzed 3348 plasma samples (6.4 TB) that were longitudinally collected from hospitalized COVID-19 patients under the auspice of the Immunophenotyping Assessment in a COVID-19 Cohort (IMPACC) study. Our parallelization strategy reduced the total runtime by ∼90% from 116 (theoretical) days to just 9 days in the high-performance computing environment. All code is open-source and can be deployed in any Simple Linux Utility for Resource Management (SLURM) high-performance computing environment, enabling the analysis of large-scale high-throughput proteomics studies.
State-of-the-art liquid chromatography/mass spectrometry (LC/MS)-based proteomic technologies, using microliter amounts of patient plasma, can detect and quantify several hundred plasma proteins in a high throughput fashion, allowing for the discovery of clinically relevant protein biomarkers and insights into the underlying pathobiological processes. Using such an in-house developed high throughput plasma proteomics allowed us to identify and quantify > 400 plasmas proteins in 15 min per sample, i.e., a throughput of 100 samples/day. We demonstrated the clinical applicability of our method in this pilot study by mapping the plasma proteomes from patients infected with human immunodeficiency virus (HIV) or herpes virus, both groups with involvement of the central nervous system (CNS). We found significant disease-specific differences in the plasma proteomes. The most notable difference was a decrease in the levels of several coagulation-associated proteins in HIV vs. herpes virus, among other dysregulated biological pathways providing insight into the differential pathophysiology of HIV compared to herpes virus infection. In a subsequent analysis, we found several plasma proteins associated with immunity and metabolism to differentiate patients with HIV-associated neurocognitive disorders (HAND) compared to cognitively normal people with HIV (PWH), suggesting the presence of plasma-based biomarkers to distinguishing HAND from cognitively normal PWH. Overall, our high-throughput plasma proteomics pipeline enables the identification of distinct proteomic signatures of HIV and herpes virus, which may help illuminate divergent pathophysiology behind virus-associated neurological disorders.
Introduction: Current techniques to diagnose and/or monitor critically ill neonates with bronchopulmonary dysplasia (BPD) require invasive sampling of body fluids, which is suboptimal in these frail neonates. We tested our hypothesis that it is feasible to use noninvasively collected urine samples for proteomics from extremely low gestational age newborns (ELGANs) at risk for BPD to confirm previously identified proteins and biomarkers associated with BPD. Methods: We developed a robust high-throughput urine proteomics methodology that requires only 50 μL of urine. We utilized the methodology with a proof-of-concept study validating proteins previously identified in invasively collected sample types such as blood and/or tracheal aspirates on urine collected within 72 h of birth from ELGANs (gestational age [26 ± 1.2] weeks) who were admitted to a single Neonatal Intensive Care Unit (NICU), half of whom eventually developed BPD (n = 21), while the other half served as controls (n = 21). Results: Our high-throughput urine proteomics approach clearly identified several BPD-associated changes in the urine proteome recapitulating expected blood proteome changes, and several urinary proteins predicted BPD risk. Interestingly, 16 of the identified urinary proteins are known targets of drugs approved by the Food and Drug Administration. Conclusion: In addition to validating numerous proteins, previously found in invasively collected blood, tracheal aspirate, and bronchoalveolar lavage, that have been implicated in BPD pathophysiology, urine proteomics also suggested novel potential therapeutic targets. Ease of access to urine could allow for sequential proteomic evaluations for longitudinal monitoring of disease progression and impact of therapeutic intervention in future studies.
Ebola virus (EBV) disease (EVD) is a highly virulent systemic disease characterized by an aggressive systemic inflammatory response and impaired vascular and coagulation systems, often leading to uncontrolled hemorrhaging and death. In this study, the proteomes of 38 sequential plasma samples from 12 confirmed EVD patients were analyzed. Of these 12 cases, 9 patients received treatment with interferon beta 1a (IFN-β-1a), 8 survived EVD, and 4 died; 2 of these 4 fatalities had received IFN-β-1a. Our analytical strategy combined three platforms targeting different plasma subproteomes: a liquid chromatography-mass spectrometry (LC-MS)-based analysis of the classical plasma proteome, a protocol that combines the depletion of abundant plasma proteins and LC-MS to detect less abundant plasma proteins, and an antibody-based cytokine/chemokine multiplex assay. These complementary platforms provided comprehensive data on 1,000 host and viral proteins. Examination of the early plasma proteomes revealed protein signatures that differentiated between fatalities and survivors. Moreover, IFN-β-1a treatment was associated with a distinct protein signature. Next, we examined those proteins whose abundances reflected viral load measurements and the disease course: resolution or progression. Our data identified a prognostic 4-protein biomarker panel (histone H1-5, moesin, kininogen 1, and ribosomal protein L35 [RPL35]) that predicted EVD outcomes more accurately than the onset viral load. IMPORTANCE As evidenced by the 2013-2016 outbreak in West Africa, Ebola virus (EBV) disease (EVD) poses a major global health threat. In this study, we characterized the plasma proteomes of 12 individuals infected with EBV, using two different LC-MS-based proteomics platforms and an antibody-based multiplexed cytokine/chemokine assay. Clear differences were observed in the host proteome between individuals who survived and those who died, at both early and late stages of the disease. From our analysis, we derived a 4-protein prognostic biomarker panel that may help direct care. Given the ease of implementation, a panel of these 4 proteins or subsets thereof has the potential to be widely applied in an emergency setting in resource-limited regions.