Advancements in mass spectrometry (MS) technologies have significantly improved the ability to quantify proteins and analyse their modifications. However, MS-based proteomics datasets frequently encounter missing values due to a complex interplay of missing at random (MAR) and missing not at random (MNAR) mechanisms. Such missing data can result in information loss and biased outcomes in data pre-processing, as well as subsequent analyses and interpretations. Few approaches effectively address both MAR and MNAR, and those that do often necessitate manual tuning of mixture percentages between them or rely on two-group experimental designs. Therefore, we developed msBayesImpute, an innovative computational method that integrates Bayesian factorization with probabilistic dropout models. We evaluated msBayesImpute against several popular imputation methods using both simulated missing values and those generated through a dilution series experiment on samples from lung cancer patients. Our comprehensive benchmark demonstrated superior performance in reconstructing missing values, estimating normalization factors, identifying differentially expressed proteins and predicting outcomes with machine learning models across varying levels of missingness and sample sizes. Notably, msBayesImpute does not require predefined experimental designs and is scalable to large-scale studies. This versatility positions msBayesImpute as an effective and robust tool for enhancing the utility of MS datasets in biological research.
Abstract Cancer cell lines are widely used in preclinical research, yet the clinical translation of findings from cell lines remains limited. Identifying cell lines that best resemble patient tumors requires integration of molecular profiles across biologically distinct sample types. Recent advances in transcriptomic integration have demonstrated the potential of deep learning for aligning data across different sample types. However, comparable approaches for proteomic data integration remain lacking, potentially because of the prevalence of missing values in proteomic datasets. Here, we introduce ProtInt, a deep learning-based framework that integrates proteomic data from cell lines and patient tumors by combining principles from proteomic imputation and transcriptomic integration methods. We applied ProtInt to integrate label-free proteomic profiles from 771 cancer cell lines and 550 treatment-naïve tumors. ProtInt outperformed batch correction and transcriptomic integration methods in aligning cell line and tumor proteomes. Comparison of the cell line proteomes before and after integration revealed recurrent increase of proteins associated with immune reaction, cell-cell communication, and interaction with the extracellular matrix, and reduction of proteins involved in transcription, post-transcriptional processing, and mitochondrial gene expression as proteomes of cell lines were adapted to resemble tumors. These results establish ProtInt as a framework for joint analysis of proteomic datasets across distinct sample types and may facilitate the identification of cell lines best suited for clinically relevant studies.
Cancer patients frequently suffer from anemia and cancer-related pain, which can be treated by non-opioid analgesics such as diclofenac (DCF) and acetaminophen (APAP) attenuating inflammatory responses. The pro-inflammatory cytokine interleukin (IL)-6 triggers the expression of acute-phase proteins, including the iron regulator hepcidin. Using proteomics and dynamic pathway modeling, we show that DCF and APAP directly impact IL-6 signaling by enhancing the induction of the feedback-inhibitor suppressor of cytokine signaling 3 (SOCS3), reducing signal transducer and activator of transcription (STAT)3 phosphorylation, and decreasing the expression of most acute-phase proteins except for hepcidin. In primary human hepatocytes (PHHs), the impact depends on the patient-specific extent of SOCS3 induction, which is anti-correlated with hepcidin expression. Whereas, in liver cancer cells, DCF and APAP stabilize the interaction of autocrine secreted bone morphogenic protein (BMP) with its receptor, resulting in strongly amplified hepcidin expression. Our studies suggest that co-inhibition of the BMP receptor counteracts excessive hepcidin production upon treatment with pain-relieving drugs and could prevent iron-deficiency-caused anemia in liver cancer. A record of this paper’s transparent peer review process is included in the supplemental information.
Idiopathic Parkinson's disease (PD) is the second most common neurodegenerative disease after Alzheimer's disease and is determined by a combination of genetic and environmental risk factors. The to date largest genome-wide association study (GWAS) on single nucleotide polymorphisms (SNPs) by Nalls at al., 2019, reported 90 SNPs that were independently associated with PD risk. However, common SNPs account only for 16-36% of the total genetic heritability of the disease suggesting that other genetic variants play a role in PD susceptibility. One example of previously understudied genetic variants are short-tandem repeats (STR, also known as microsatellites), i.e., repeating sequence motifs in the human genome of 1-6 nucleotides in length. Thus, in this study, we performed a GWAS on imputed STRs in a large PD case-control dataset (n=4,757), and meta-analyzed these data with those from a previous study of the International PD Genetics Consortium (Bustos et al., 2023) resulting in a total sample size of 43,844 PD cases and controls. Thus, in this study, we performed a GWAS on imputed STRs in a large PD case-control dataset (n=4,757) from the US (the PEG and a GHC-based study) and Denmark (PASIDA).
Prerequisite for a successful proteomics experiment is a high-quality lysis of the sample of interest, resulting in a large number of identified proteins as well as a high coverage of protein sequences. Therefore, the choice of suitable lysis conditions is crucial. Many buffers were previously employed in proteomics studies, yet a comprehensive comparison of lysate preparation conditions was so far missing. In this study, we compared the efficiency of four commonly used lysis buffers, containing the agents NP40, SDS, urea or GdnHCl, in four different types of biological samples (suspension and adherent cell lines, primary mouse cells and mouse liver tissue). After liquid chromatography-mass spectrometry (LC-MS) measurement and MaxQuant analysis, we compared chromatograms, intensities, number of identified proteins and the localization of the identified proteins. Overall, SDS emerged as the most reliable reagent, ensuring stable performance and reproducibility across diverse samples. Furthermore, our data advocated for a dual-sample lysis approach, including that the resulting pellet is lysed again after the initial lysis with a urea lysis buffer and subsequently both lysates are combined for a single LC-MS run to maximize the proteome coverage. However, none of the investigated lysis buffers proved to be superior in every category, indicating that the lysis buffer of choice depends on the proteins of interest and on the biological question. Further, we demonstrated with our systematic studies the establishment of conditions that allows to perform global proteomics and affinity purification-based interactome characterization from the same lysate. In sum our results provide guidance for the best-suited lysis buffer for mass spectrometry-based proteomics depending on the question of interest.### Competing Interest StatementThe authors have declared no competing interest.
Chronic liver diseases are worldwide on the rise. Due to the rapidly increasing incidence, in particular in Western countries, metabolic dysfunction-associated steatotic liver disease (MASLD) is gaining importance as the disease can develop into hepatocellular carcinoma. Lipid accumulation in hepatocytes has been identified as the characteristic structural change in MASLD development, but molecular mechanisms responsible for disease progression remained unresolved. Here, we uncover in primary hepatocytes from a preclinical model fed with a Western diet (WD) an increased basal MET phosphorylation and a strong downregulation of the PI3K-AKT pathway. Dynamic pathway modeling of hepatocyte growth factor (HGF) signal transduction combined with global proteomics identifies that an elevated basal MET phosphorylation rate is the main driver of altered signaling leading to increased proliferation of WD-hepatocytes. Model-adaptation to patient-derived hepatocytes reveal patient-specific variability in basal MET phosphorylation, which correlates with patient outcome after liver surgery. Thus, dysregulated basal MET phosphorylation could be an indicator for the health status of the liver and thereby inform on the risk of a patient to suffer from liver failure after surgery.
File 3 L1236 Data from Dynamic Mathematical Modeling of IL13-Induced Signaling in Hodgkin and Primary Mediastinal B-Cell Lymphoma Allows Prediction of Therapeutic Targets
Cancer is a devastating disease and the second leading cause of death worldwide. However, the development of resistance to current therapies is making cancer treatment more difficult. Combining the multi-omics data of individual tumors with information on their in-vitro Drug Sensitivity and Resistance Test (DSRT) can help to determine the appropriate therapy for each patient. Miniaturized high-throughput technologies, such as the droplet microarray, enable personalized oncology. We are developing a platform that incorporates DSRT profiling workflows from minute amounts of cellular material and reagents. Experimental results often rely on image-based readout techniques, where images are often constructed in grid-like structures with heterogeneous image processing targets. However, manual image analysis is time-consuming, not reproducible, and impossible for high-throughput experiments due to the amount of data generated. Therefore, automated image processing solutions are an essential component of a screening platform for personalized oncology. We present our comprehensive concept that considers assisted image annotation, algorithms for image processing of grid-like high-throughput experiments, and enhanced learning processes. In addition, the concept includes the deployment of processing pipelines. Details of the computation and implementation are presented. In particular, we outline solutions for linking automated image processing for personalized oncology with high-performance computing. Finally, we demonstrate the advantages of our proposal, using image data from heterogeneous practical experiments and challenges.
In biomedical engineering, deep neural networks are commonly used for the diagnosis and assessment of diseases through the interpretation of medical images. The effectiveness of these networks relies heavily on the availability of annotated datasets for training. However, obtaining noise-free and consistent annotations from experts, such as pathologists, radiologists, and biologists, remains a significant challenge. One common task in clinical practice and biological imaging applications is instance segmentation. Though, there is currently a lack of methods and open-source tools for the automated inspection of biomedical instance segmentation datasets concerning noisy annotations. To address this issue, we propose a novel deep learning-based approach for inspecting noisy annotations and provide an accompanying software implementation, AI 2 Seg, to facilitate its use by domain experts. The performance of the proposed algorithm is demonstrated on the medical MoNuSeg dataset and the biological LIVECell dataset.
L1236 model from Dynamic Mathematical Modeling of IL13-Induced Signaling in Hodgkin and Primary Mediastinal B-Cell Lymphoma Allows Prediction of Therapeutic Targets
The neuron-glia cross-talk is critical to brain homeostasis and is particularly affected by neurodegenerative diseases. How neurons manipulate the neuron-astrocyte interaction under pathological conditions, such as hyperphosphorylated tau, a pathological hallmark in Alzheimer’s disease (AD), remains elusive. In this study, we identified excessively elevated neuronal expression of adenosine receptor 1 (Adora1 or A1R) in 3×Tg mice, MAPT P301L (rTg4510) mice, patients with AD, and patient-derived neurons. The up-regulation of A1R was found to be tau pathology dependent and posttranscriptionally regulated by Mef2c via miR-133a-3p. Rebuilding the miR-133a-3p/A1R signal effectively rescued synaptic and memory impairments in AD mice. Furthermore, neuronal A1R promoted the release of lipocalin 2 (Lcn2) and resulted in astrocyte activation. Last, silencing neuronal Lcn2 in AD mice ameliorated astrocyte activation and restored synaptic plasticity and learning/memory. Our findings reveal that the tau pathology remodels neuron-glial cross-talk and promotes neurodegenerative progression. Approaches targeting A1R and modulating this signaling pathway might be a potential therapeutic strategy for AD.
Supplementary Methods, Figures and References from Dynamic Mathematical Modeling of IL13-Induced Signaling in Hodgkin and Primary Mediastinal B-Cell Lymphoma Allows Prediction of Therapeutic Targets
To address the challenge of drug resistance and limited treatment options for recurrent gliomas with IDH1 mutations, a highly miniaturized screening of 2208 FDA-approved drugs is conducted using a high-throughput droplet microarray (DMA) platform. Two patient-derived temozolomide-resistant tumorspheres harboring endogenous IDH1 mutations (IDH1mut ) are utilized. Screening identifies over 20 drugs, including verteporfin (VP), that significantly affected tumorsphere formation and viability. Proteomics analysis reveals that nuclear pore complex may be a potential VP target, suggesting a new mechanism of action independent of its known effects on YAP1. Knockdown experiments exclude YAP1 as a drug target in tumorspheres. Pathway analysis shows that NUP107 is a potential upstream regulator associated with VP response. Analysis of publicly available genomic datasets shows a significant correlation between high NUP107 expression and decreased survival in IDH1mut astrocytoma, suggesting NUP107 may be a potential biomarker for VP response. This study demonstrates a miniaturized approach for cost-effective drug repurposing using 3D glioma models and identifies nuclear pore complex as a potential target for drug development. The findings provide preclinical evidence to support in vivo and clinical studies of VP and other identified compounds to treat IDH1mut gliomas, which may ultimately improve clinical outcomes for patients with this challenging disease.
Background: Macrophages play an important role in maintaining liver homeostasis and regeneration. However, it is not clear to what extent the different macrophage populations of the liver differ in terms of their activation state and which other liver cell populations may play a role in regulating the same. Methods: Reverse transcription PCR, flow cytometry, transcriptome, proteome, secretome, single cell analysis, and immunohistochemical methods were used to study changes in gene expression as well as the activation state of macrophages in vitro and in vivo under homeostatic conditions and after partial hepatectomy. Results: We show that F4/80 + /CD11b hi /CD14 hi macrophages of the liver are recruited in a C-C motif chemokine receptor (CCR2)–dependent manner and exhibit an activation state that differs substantially from that of the other liver macrophage populations, which can be distinguished on the basis of CD11b and CD14 expressions. Thereby, primary hepatocytes are capable of creating an environment in vitro that elicits the same specific activation state in bone marrow–derived macrophages as observed in F4/80 + /CD11b hi /CD14 hi liver macrophages in vivo . Subsequent analyses, including studies in mice with a myeloid cell–specific deletion of the TGF-β type II receptor, suggest that the availability of activated TGF-β and its downregulation by a hepatocyte-conditioned milieu are critical. Reduction of TGF-βRII-mediated signal transduction in myeloid cells leads to upregulation of IL-6, IL-10, and SIGLEC1 expression, a hallmark of the activation state of F4/80 + /CD11b hi /CD14 hi macrophages, and enhances liver regeneration. Conclusions: The availability of activated TGF-β determines the activation state of specific macrophage populations in the liver, and the observed rapid transient activation of TGF-β may represent an important regulatory mechanism in the early phase of liver regeneration in this context.
Due to the broad use of deep learning and its need for big data, annotated and available databases for different tasks are constantly appearing. Nevertheless, they often remain unexploited due to the difficulty of effectively performing transfer learning between different databases. In medical imaging, the task of transfer learning is challenging due to: the variety of image modalities, organ/cell shapes, etc., and the lack of available and annotated data. In this paper, we propose an automated pipeline for predicting the similarity values of new database compared to known annotated databases. The system consists of an autoencoder trained on a comprehensive loss function that considers image reconstruction, style features, and dataset membership. A similarity measure is defined based on the resulting 2D latent space, which is demonstrated to have a correlation with the pre-training results on not annotated databases. Hence, our similarity measure could be used to select the most suitable known database for transfer learning or domain adaptation.
DNA methylation (DNAm) is an epigenetic mark with essential roles in disease development and predisposition. Here, we created genome-wide maps of methylation quantitative trait loci (meQTL) in three peripheral tissues and used Mendelian randomization (MR) analyses to assess the potential causal relationships between DNAm and risk for two common neurodegenerative disorders, i.e. Alzheimer's disease (AD) and Parkinson's disease (PD). Genome-wide single nucleotide polymorphism (SNP; ~5.5M sites) and DNAm (~850K CpG sites) data were generated from whole blood (n=1,058), buccal (n=1,527) and saliva (n=837) specimens. We identified between 11 and 15 million genome-wide significant (p<10-14) SNP-CpG associations in each tissue. Combining these meQTL GWAS results with recent AD/PD GWAS summary statistics by MR strongly suggests that the previously described associations between PSMC3, PICALM, and TSPAN14 and AD may be founded on differential DNAm in or near these genes. In addition, there is strong, albeit less unequivocal, support for causal links between DNAm at PRDM7 in AD as well as at KANSL1/MAPT in AD and PD. Our study adds valuable insights on AD/PD pathogenesis by combining two high-resolution "omics" domains, and the meQTL data shared along with this publication will allow like-minded analyses in other diseases.
Background Studies on DNA methylation (DNAm) in Alzheimer’s disease (AD) have recently highlighted several genomic loci showing association with disease onset and progression. Methods Here, we conducted an epigenome-wide association study (EWAS) using DNAm profiles in entorhinal cortex (EC) from 149 AD patients and control brains and combined these with two previously published EC datasets by meta-analysis (total n = 337). Results We identified 12 cytosine-phosphate-guanine (CpG) sites showing epigenome-wide significant association with either case–control status or Braak’s tau-staging. Four of these CpGs, located in proximity to CNFN/LIPE , TENT5A, PALD1/PRF1, and DIRAS1 , represent novel findings. Integrating DNAm levels with RNA sequencing-based mRNA expression data generated in the same individuals showed significant DNAm-mRNA correlations for 6 of the 12 significant CpGs. Lastly, by calculating rates of epigenetic age acceleration using two recently proposed “epigenetic clock” estimators we found a significant association with accelerated epigenetic aging in the brains of AD patients vs. controls. Conclusion In summary, our study represents the hitherto most comprehensive EWAS in AD using EC and highlights several novel differentially methylated loci with potential effects on gene expression.
RNA abundance is tightly regulated in eukaryotic cells by modulating the kinetic rates of RNA production, processing, and degradation. To date, little is known about time‐dependent kinetic rates during dynamic processes. Here, we present SLAM‐Drop‐seq, a method that combines RNA metabolic labeling and alkylation of modified nucleotides in methanol‐fixed cells with droplet‐based sequencing to detect newly synthesized and preexisting mRNAs in single cells. As a first application, we sequenced 7280 HEK293 cells and calculated gene‐specific kinetic rates during the cell cycle using the novel package Eskrate. Of the 377 robust‐cycling genes that we identified, only a minor fraction is regulated solely by either dynamic transcription or degradation (6 and 4%, respectively). By contrast, the vast majority (89%) exhibit dynamically regulated transcription and degradation rates during the cell cycle. Our study thus shows that temporally regulated mRNA degradation is fundamental for the correct expression of a majority of cycling genes. SLAM‐Drop‐seq, combined with Eskrate, is a powerful approach to understanding the underlying mRNA kinetics of single‐cell gene expression dynamics in continuous biological processes.
Objective : The scarcity of high-quality annotated data is omnipresent in machine learning. Especially in biomedical segmentation applications, experts need to spend a lot of their time into annotating due to the complexity. Hence, methods to reduce such efforts are desired. Methods : Self-Supervised Learning (SSL) is an emerging field that increases performance when unannotated data is present. However, profound studies regarding segmentation tasks and small datasets are still absent. A comprehensive qualitative and quantitative evaluation is conducted, examining SSL's applicability with a focus on biomedical imaging. We consider various metrics and introduce multiple novel application-specific measures. All metrics and state-of-the-art methods are provided in a directly applicable software package ( https://osf.io/gu2t8/ ). Results : We show that SSL can lead to performance improvements of up to 10%, which is especially notable for methods designed for segmentation tasks. Conclusion : SSL is a sensible approach to data-efficient learning, especially for biomedical applications, where generating annotations requires much effort. Additionally, our extensive evaluation pipeline is vital since there are significant differences between the various approaches. Significance : We provide biomedical practitioners with an overview of innovative data-efficient solutions and a novel toolbox for their own application of new approaches. Our pipeline for analyzing SSL methods is provided as a ready-to-use software package.
Background Dysregulation of microRNA (miRNA)-mediated gene expression has been implicated in the pathogenesis and course of many neurodegenerative diseases including Parkinson’s disease (PD). However, the functionally relevant miRNAs remain largely unknown. Previous meta-analyses on differential miRNA expression data in post-mortem PD brains have highlighted several miRNAs showing consistent and statistically significant effects. However, these meta-analyses were based on exceedingly small sample sizes. Methods In this study, we quantified the expression of the four most compelling PD candidate miRNAs from these meta-analyses in the superior temporal gyrus (STG) of one of the largest case-control post-mortem brain datasets available (261 samples), thereby quadruplicating previously investigated sample sizes. Furthermore, we probed for common differential miRNA expression signatures with Alzheimer’s disease (AD) by also analyzing these miRNAs in post-mortem STG of 190 AD patients and controls and by testing six top AD miRNAs in the PD brains. Results Of all ten analyzed miRNAs, PD candidate miRNA homo sapiens (hsa-) miR-132-3p showed evidence for differential expression in both PD (p=4.89E-06) and AD (p=3.20E-24), and AD miRNAs hsa-miR-132-5p (p=4.52E-06) and hsa-miR-129-5p (p=0.0379) showed evidence for differential expression in PD. Combining these novel data with previously published data substantially improved the statistical support (α=3.85E-03 using Bonferroni correction) of the corresponding meta-analyses clearly and compellingly implicating these miRNAs in both PD and AD. Furthermore, hsa-miR-132-3p/-5p (but not hsa-miR-129-5p) showed association with neuropathological Braak PD staging (p=3.51E-03/p=0.0117), suggesting that these miRNAs may play a role in α-synuclein aggregation beyond the early disease phase. Conclusions Our study represents the largest independent assessment of recently highlighted candidate brain miRNAs in PD and AD post-mortem brain samples, to date. Our results implicate hsa-miR-132-3p/-5p and hsa-miR-129-5p to be differentially expressed in both PD and AD brains, potentially pinpointing shared pathogenic mechanisms across these neurodegenerative diseases.