Proteomic and phenotypic cell sensitivity datasets are increasingly important for understanding chemoproteomics and the underlying drug mechanisms of action. Yet, integrating such heterogeneous datasets remains challenging due to inconsistent annotations, incompatible IDs, and variable data processing methods. Here, a major update to ProteomicsDB (https://www.proteomicsdb.org) is presented that combines over 1300 proteomic and 1000 transcriptomic profiles with phenotypic cell sensitivity data across >1500 human cancer cell lines and 1470 drugs. Harmonizing cell line and drug names and applying a standardized normalization and refitting pipeline for dose–response curves enables consistent, statistically robust analysis across studies. Three new graphical user interfaces support interactive exploration of cell sensitivity data, exploring the protein targets and dose-resolved changes in protein expression in the presence of a drug, and comparing the expression profiles of cell lines. With this update, ProteomicsDB is strengthening its future role as a central hub for proteomics and multi-omics, providing researchers with a unified framework to explore phenotypic cell sensitivity in combination with dose-resolved expression proteomics at the molecular level, supporting biomarker discovery, drug repurposing, and precision medicine applications.
Advancements in spatial transcriptomics and single-cell RNA sequencing have enhanced our understanding of gene expression within tissues. Spatial transcriptomics retains spatial context at the expense of resolution, often resulting in cell mixtures, whereas single-cell RNA sequencing offers single-cell resolution with the loss of spatial information. Some computational methods aim to integrate data from these two technologies; however, a ground truth for their evaluation is typically lacking. Thus, simulation techniques may be used to generate artificial gold or silver standards, offering the possibility for standardized analysis. This not only requires an accurate replication of real tissue types, but also sufficient sample diversity, calling for a unified evaluation of these properties between present techniques. Existing benchmark metrics and platforms often favor simulations that closely replicate the input data rather than promoting novel tissue layouts. This paper introduces a comprehensive benchmarking platform that evaluates spatial transcriptomics simulation methods across data property distributions, biological signal preservation, and similarity-based metrics. Our framework ensures that simulations go beyond simple data replication, instead introducing biologically meaningful variation. BEASTsim can be easily integrated into analysis pipelines and provides a practical tool for evaluating and developing computational methods, thereby advancing the integration of spatial transcriptomics and single-cell RNA sequencing data to yield more accurate biological insights. As a result, we have utilized BEASTsim to create a decision tree that helps users select the most suitable simulation model based on their data and goals. This work provides a practical tool for evaluating and developing computational methods, thereby advancing the integration of spatial transcriptomics and single-cell RNA sequencing data to yield more accurate biological insights.
Resting metabolic rate (RMR) is modulated by a variety of factors. Accurate prediction of RMR is essential for planning energy requirements but remains challenging due to interindividual variability. This study aimed to develop and evaluate machine learning models for predicting RMR using comprehensive data from the cross-sectional enable study and to identify the most predictive and stable features across different study populations. RMR was predicted using data from 454 participants of the enable phenotyping platform (Freising and Nuremberg cohort). We systematically compared linear and nonlinear machine learning models trained on either the full set of 94 predictors or a reduced set of routinely accessible variables, including sex, age, body weight, fat mass, and fat-free mass. Model performance was assessed by cross-validation. The best-performing model (Lasso) was further evaluated on independent test datasets from other cohorts. Feature importance and stability were assessed using repeated cross-validation and marginal variance decomposition. Lasso regression consistently outperformed other models, particularly when trained on the enable cohort feature set. The final model explained 76.8% of RMR variance in the Freising cohort. Key predictive features included fat-free mass, body weight, and mean outdoor temperature. Blood-based features contributed marginally, whereas microbiota and fecal short-chain fatty acids variables did not contribute to explaining RMR. This novel prediction model for RMR shows improved accuracy in comparison with traditional models. Although microbiota composition did not contribute to explain the residual variation in RMR, the inclusion of clinical blood parameters and outdoor temperature improved predictive performance. Clinical Trial Registry Number: DRKS00009797.NEW & NOTEWORTHY We introduce a novel machine learning framework for predicting resting metabolic rate (RMR), emphasizing the superior performance of Lasso regression. Our analysis incorporates both standard clinical variables and previously underexplored factors such as gut microbiota, fecal short-chain fatty acids (SCFAs), and mean outdoor temperature.
In silico cell-type deconvolution from bulk transcriptomics data is a powerful technique to gain insights into the cellular composition of complex tissues. While first-generation methods used precomputed expression signatures covering limited cell types and tissues, second-generation tools use single-cell RNA sequencing data to build custom signatures for deconvoluting arbitrary cell types, tissues, and organisms. This flexibility poses significant challenges in assessing their deconvolution performance. Here, we comprehensively benchmark second-generation tools, disentangling different sources of variation and bias using a diverse panel of real and simulated data. Our results reveal substantial differences in accuracy, scalability, and robustness across methods, depending on factors such as cell-type similarity, reference composition, and dataset origin. Our study highlights the strengths, limitations, and complementarity of state-of-the-art tools, shedding light on how different data characteristics and confounders impact deconvolution performance. We provide the scientific community with an ecosystem of tools and resources, omnideconv, simplifying the application, benchmarking, and optimization of deconvolution methods.
Focal segmental glomerulosclerosis (FSGS) is a major cause of nephrotic syndrome and progression to end-stage renal disease, yet its molecular pathogenesis remains still incompletely defined. While transcriptional alterations in podocytes have been extensively characterized, the contribution of post-transcriptional regulatory mechanisms is poorly understood. Here, we combined a zebrafish podocyte-specific injury model with glomerulus-resolved transcriptomic profiling to dissect RNA regulatory alterations during FSGS progression. Integrated analyses of bulk RNA sequencing, small RNA profiling, and alternative splicing revealed pronounced, time-dependent remodeling of the glomerular transcriptome. We demonstrate that podocyte injury is associated with loss of key podocyte-specific proteins, activation of inflammatory pathways, remodeling of the extracellular matrix, and altered microRNA expression, such as miR-21 and miR-193. Moreover, we found that alternative splicing influences key podocyte gene expression, affecting genes critical for slit diaphragm integrity, actin cytoskeleton organization, and glomerular basement membrane stability. Isoform analyses identified FSGS-associated isoform switches in SRSF3 and EPB41L5. Importantly, these changes were also evident in glomeruli from FSGS patients, demonstrating that the zebrafish model recapitulates key molecular features of human disease and highlighting alternative splicing as a central regulatory mechanism in FSGS. Post-transcriptional regulations, such as alternative splicing and microRNA dysregulation, are identified as central and underappreciated processes in injured podocytes in focal segmental glomerulosclerosis (FSGS), with disease-associated isoform switches in SRSF3 and EPB41L5. Post-transcriptional regulations, such as alternative splicing and microRNA dysregulation, are identified as central and underappreciated processes in injured podocytes in focal segmental glomerulosclerosis (FSGS), with disease-associated isoform switches in SRSF3 and EPB41L5.
Circular RNAs have garnered considerable interest, as they have been implicated in numerous biological processes and diseases. Through their stability, they are often considered promising biomarker candidates or therapeutic targets. Due to the lack of a poly(A) tail, circRNAs are best detected in total RNA-seq data after depleting ribosomal RNA. However, we observe that the application of circRNA detection in the vastly more ubiquitous poly(A)-enriched RNA-seq data still occurs. In this study, we systematically compare the detection of circRNAs in two matched poly(A) and ribosomal RNA-depleted data sets. Our results indicate that the comparably few circRNAs detected in poly(A) data are likely false positives. In addition, we demonstrate that the quality of sample processing, as measured by the fraction of ribosomal reads, significantly affects the sensitivity of circRNA detection, leading to a bias in downstream analysis. Our findings establish best practices for circRNA research: total RNA sequencing with effective rRNA depletion is the preferred approach for accurate circRNA profiling, whereas poly(A)-enriched data are unsuitable for comprehensive detection. Employing multiple circRNA detection tools and prioritizing back-splice junctions identified by several algorithms enhances confidence in the selection of candidates. These recommendations, validated across diverse datasets and tissue types, provide generalizable principles for robust circRNA analysis. ### Competing Interest Statement The authors have declared no competing interest. Funded by the Federal Ministry of Education and Research (BMBF) and the Free State of Bavaria under the Excellence Strategy of the Federal Government and the Länder, as well as by the Technical University of Munich Institute for Advanced Study, Garching, Germany through an Anna Boyksen Fellowship (P.A.F.)
Protein-protein interaction (PPI) databases do not faithfully reflect biological realities. Instead, they are influenced by study and technical biases that distort certain protein and interaction attributes. Machine learning models can exploit these as learning shortcuts if the negative dataset is not constructed with care. So far, the shortcuts introduced during PPI dataset construction have only been examined in isolation. Here, we systematically characterize both reported and, to our knowledge, previously unreported biases in PPI datasets that lead machine learning models to learn shortcuts instead of biological signal. We analyze HIPPIE, IntAct, and STRING, dedicated PPI databases, as well as two datasets derived from 3D-structural information in the Protein Data Bank (PDB). We show that random data splitting introduces strong topological shortcuts. When train-test protein overlap is removed, the resulting datasets still retain usable shortcuts stemming from self-interactions, taxonomic identity, and functional relatedness, whose prevalence interestingly depends on the data source. We further show that sampling negatives from a set of high-confidence non-interactors, an intuitively appealing choice, can amplify the shortcut stemming from functional relatedness. To detect and mitigate these biases, we provide an open Nextflow pipeline that combines similarity-aware, data-loss-minimizing dataset splitting with bias-minimizing negative sampling, both formulated as integer linear programs. Its key concept of quantifying biases to minimize them through optimization-based negative sampling can, in principle, be extended to any machine learning problem where the pool of negative candidates is much larger than the positives and is thus of interest also beyond PPI prediction.
BACKGROUND AND PURPOSE:Complex diseases often lack an actionable understanding of their underlying causal biological mechanisms, which leads to treating symptoms rather than causes. Network and systems medicine define disease mechanisms through disease-associated genes, their encoded proteins and their protein-protein interactions (PPIs), thus forming disease modules. Complex diseases can be subdivided into actionable causal mechanisms for potential precision and curative therapy by repurposing small-molecule drugs for new indications. However, current computational methods for disease module construction overlook pathway annotations, cellular compartments and directed PPIs. Consequently, disease modules require contextual refinement to identify dysregulations, select appropriate drug classes and eliminate promiscuous proteins. EXPERIMENTAL APPROACH:Here, we present Drugst.One DREAM, which equips biomedical experts with a user-friendly toolbox for disease module refinement that does not require bioinformatics expertise. This extension of the web tool Drugst.One introduces network editing features. Users can refine PPI modules supported by pathway enrichment analysis and network clustering. Dedicated graph layouts highlight subcellular localisation and causal relationships queried from OmniPath. KEY RESULTS:We demonstrate our tool by reproducing a previously described NOX5-containing module and refining an algorithmically inferred candidate module for Crohn's disease, showcasing its effectiveness in refining disease modules for a broad user group in pharmacology and biomedical research. CONCLUSION AND IMPLICATIONS:The Drugst.One DREAM extension closes an important gap in the network medicine tool landscape by offering experts a user-friendly option for refining disease modules.
The sequence of the human genome provides a foundation for understanding cellular processes in health and disease1. The organisation of this primary genetic information into cell-specific structure and function is critical to understanding the cell type-specific interpretation and execution of the genome. Epigenetic processes are essential for packaging and higher-level functional organisation of the genome, and changes therein are increasingly recognised as contributors to human disease. Building on primary data generated by multinational consortia, the International Human Epigenome Consortium2 (IHEC) has uniformly processed a collection of more than 2000 comprehensive human reference epigenomes, collectively referred to as EpiATLAS. This effort involved the development of standardised molecular and bioinformatics protocols, metadata models, and analytical tools to manage, integrate, display, and share vast amounts of epigenomic data. This includes the creation of a publicly available Epigenome Reference Registry, which provides a system for accessing protected human subject datasets and facilitates open searching of de-identified samples and experimental data. The integrated EpiATLAS ecosystem and its comprehensive human reference epigenome maps provide an unprecedented resource for the biosciences, expanding the annotated epigenomic landscape while uncovering previously unappreciated relationships among regulatory layers and revealing how epigenetic inputs underpin fundamental cellular functions and disease associations.
Motivation Most diseases result from complex molecular interactions of genes and proteins. Various network-based methods characterize these mechanisms by expanding seed genes into disease modules. Their underlying algorithmic strategies differ, making it difficult to determine which of the created modules are most useful or biologically plausible.Results To address this challenge, we developed an all-in-one pipeline that handles installation, input preparation, execution, and systematic evaluation of six widely used module detection tools, considering module topology, functional coherence, robustness, and the capacity to recover seeds. To showcase the value of our pipeline and provide guidance to potential users, we conducted a comprehensive evaluation across 50 different disease-network combinations, revealing substantial variability among the derived disease modules, driven by both network and algorithm choices. We show that methods are robust to minor perturbations but struggle to recover omitted seeds. None consistently outperforms all others, underscoring the need for careful method selection. Our work enables the systematic comparison of disease module discovery approaches and promotes reproducible network medicine research. Integrated into the nf-core project, it is intended as an extendable, long-term resource for tracking progress in the field.Availability and Implementation The pipeline is implemented in Nextflow. Code and documentation are available through GitHub (https://github.com/nf-core/diseasemodulediscovery) and the nf-core website (https://nf-co.re/diseasemodulediscovery). Code and data used for demonstrating the pipeline are available through GitHub (https://github.com/REPO4EU/modulediscovery_demonstration).
Abstract Antihormonal therapies such as selective oestrogen receptor modulators like tamoxifen or aromatase inhibitors like letrozole represent a cornerstone for breast cancer prevention and therapy of oestrogen receptor-positive breast cancer. Therapeutic monitoring can include blood tests and imaging; however, genetically-based approaches are not yet in practice. Ideally, a test would be able to detect a positive molecular response across different oestrogen pathway-suppressive approaches. Circular RNAs are a species of non-coding RNAs detectable in plasma that have been proposed as non-invasive therapeutic biomarkers. To determine whether a set of specific circular RNAs is altered across oestrogen-suppressive pathway approaches, we analysed mammary gland-specific total RNA sequencing data from two individual genetically engineered mouse models (GEMMs) of oestrogen pathway-induced breast cancer, with or without exposure to tamoxifen or letrozole. The nf-core/circrna pipeline was used to identify candidate circRNA regions that were differentially expressed in response to either tamoxifen or letrozole. We then screened for candidate circRNA regions that were differentially regulated by both antihormonals. Four up-regulated and 31 down-regulated candidate circRNA regions with host genes known to be expressed in human breast epithelial cells were identified as showing reproducible differential regulation in response to antihormonal treatment.
Anti-hormonal therapies such as selective estrogen receptor modulators like tamoxifen or aromatase inhibitors like letrozole represent a cornerstone for breast cancer prevention and therapy of estrogen receptor-positive breast cancer. Therapeutic monitoring can include blood tests and imaging; however, genetically-based approaches are not yet in practice. Ideally, a test would be able to detect a positive molecular response across different estrogen pathway-suppressive approaches. Circular RNAs are a species of non-coding RNAs detectable in plasma that have been proposed as non-invasive therapeutic biomarkers. To determine whether a set of specific circular RNAs is altered across estrogen-suppressive pathway approaches, we analyzed mammary gland-specific total RNA sequencing data from two individual genetically engineered mouse models (GEMMs) of estrogen pathway-induced breast cancer, with or without exposure to tamoxifen or letrozole. The nf-core/circrna pipeline was used to identify circRNAs that were differentially expressed in response to either tamoxifen or letrozole. We then screened for circRNAs that were differentially regulated by both anti-hormonals. Four up-regulated and 31 down-regulated circRNAs with host genes known to be expressed in human breast epithelial cells were identified as showing reproducible differential regulation in response to anti-hormonal treatment.
The chromatin organizer SATB1 is indispensable for thymic regulatory T cell (Treg cell) development and T helper cell induction. Several gene loci have been described to be SATB1-controlled, including the transcription factor GATA3 and the cytokine loci IL-4 and IL-17. However, the global effects of SATB1 on fully differentiated human CD4 conventional T cells (Tconv cells) and Treg cells, and thus the potential of SATB1 as a target for T-cell engineering, are poorly understood. Here, we describe SATB1-regulated gene signatures as largely subset-specific, with broader effects on Treg cells. Despite distinct gene-regulatory patterns, we observe overarching dysregulated cytokine and JAK-STAT signaling after SATB1 ablation. Functionally, SATB1 KO reduces suppressive capacities of human Treg cells but boosts tumor clearance via CD4 CAR T cells in a preclinical, humanized mouse model. Taken together, Treg destabilization and simultaneous increased activation of CD4 CAR T cells by SATB1 modulation may be a strategy to boost the efficiency of CAR T cell therapies.
Motivation: A growing volume of large-scale genome-wide association study (GWAS) datasets offers unprecedented power to uncover the genetic determinants of complex traits, but existing web-based platforms for GWAS data exploration provide limited support for interpreting these findings within broader biological systems. Systems medicine is particularly well-suited to fill this gap, as its network-oriented view of molecular interactions enables the integration of genetic signals into coherent network modules, thereby opening opportunities for disease mechanism mining and drug repurposing. Results: We introduce GNExT (GWAS network exploration tool), a web-based platform that moves beyond the variant-level effect and significance exploration provided by existing solutions. By including MAGMA and Drugst.One, GNExT allows its users to study genetic variants on the network level down to the identification of potential drug repurposing candidates. Moreover, GNExT advances over the current state of the art by offering a highly standardized Nextflow pipeline for data import and preprocessing, allowing researchers to easily deploy their study results on a web interface. We demonstrate the utility of GNExT using a genome-wide association meta-analysis of human olfactory identification, in which the framework translated isolated GWAS signals to potential pharmacological targets in human olfaction. Availability and Implementation: The complete GNExT ecosystem, including the Nextflow preprocessing pipeline, the backend service, and frontend interface, is publicly available on GitHub (https://github.com/dyhealthnet/gnext\_nf\_pipeline, https://github.com/dyhealthnet/gnext_platform). The public instance of the GNExT platform on olfaction is available under http://olfaction.gnext.gm.eurac.edu. ### Competing Interest Statement M.L. consults for mbiomics GmbH. All other authors declare no competing interest. Deutsche Forschungsgemeinschaft, 516188180 Autonomous Province of Bolzano-Bozen, Joint Projects South TyrolGermany 2024 FRQS, chercheur boursier sénior 352197 NSERC, RGPIN-2022-04813
BACKGROUND:Inflammatory bowel disease (IBD), including ulcerative colitis (UC) and Crohn's disease (CD), is associated with changes in the gut microbiome. Studies comparing fecal, gut mucosal, and salivary microbiomes are rare, and questions remain regarding the interaction of these compartments. METHODS:In this case-control study, 16S rRNA gene amplicon sequencing was performed on samples from stool, intestinal mucosa, and saliva of 120 patients with IBD. Patients with signs of non-IBD colonic inflammation (N = 28) and healthy subjects (N = 67) served as controls. A total of 480 16S profiles were analyzed. The results were evaluated with multiple clinical and pathological parameters and potential confounders were considered. The study aimed to find microbial biomarkers specific to IBD and signatures of intestinal barrier dysfunction. RESULTS:Fecal α-diversity of IBD patients was reduced and Pseudomonas species was significantly increased in the mucosa of IBD patients (Pseudomonas-positive mucosa [PSM positive], P value < .001, Mann-Whitney U test). Comparison of matched stool and mucosa samples showed high abundance of Pseudomonas species in gut mucosa but not in fecal samples, especially in CD patients. Interestingly, in PSM positive, Paracoccus species, Bacteroides species, and Streptococcus species were more abundant. Importantly, the results were independent of disease severity, histopathology, medication, and other metadata. CONCLUSIONS:The opportunistic pathogenic bacterium Pseudomonas species is more prevalent in the gut mucosa of patients with IBD. This indicates a disruption of the gut barrier with increasing mucosal colonization or invasion of the bacteria. The finding is independent of clinical metadata and confounders and occurs in new-onset IBD but not in non-IBD intestinal inflammation, which suggests disease specificity.
Large-scale drug sensitivity screens have enabled training drug response prediction models based on cancer cell line omics profiles to advance personalized medicine. While model performances reported in the literature appear promising, successful translation to the clinic remains limited. In this work, we discuss key obstacles that lead to overly optimistic performance estimates of state-of-the-art models, making it challenging to track progress in the field. To address them, we present DrEval, a pipeline for unbiased, biologically meaningful evaluation of cancer drug response models. DrEval is designed as a living open-source benchmark that integrates baseline and literature models with standardized hyperparameter tuning, statistically rigorous evaluation, cross-study benchmarks, and supports ablation studies and publication-ready visualizations. Using DrEval, we show that deep learning models barely outperform a naive model that predicts only the mean drug and cell line effects, while no complex model outperforms properly tuned tree-based ensemble baselines in relevant settings.
Direct cellular reprogramming, converting one differentiated cell type directly into another, holds immense promise for regenerative medicine, developmental biology, and disease modeling. Identifying optimal transcription factor (TF) combinations to control this process remains complex and labor-intensive. Over the last decade, various computational tools emerged to infer TF sets for reprogramming. However, current methodologies possess critical limitations, and the absence of robust benchmarking standards makes it impossible to precisely validate and compare their performance. To address these challenges, we present a comprehensive analysis of existing computational methods for direct reprogramming and introduce a web application designed to support researchers in identifying and validating optimal TF sets. Our platform integrates predictions from established tools, incorporates a state-of-the-art Retrieval-Augmented Generation (RAG) system for efficient literature querying, and offers tools to further validate predictions. By providing a unified and interactive resource, our web application enhances the accessibility and efficiency of TF discovery for direct reprogramming. Furthermore, we discuss critical limitations shared by current methodologies and highlight the need for computational tools that can account for the complex regulatory dynamics of direct reprogramming. This work not only advances the toolkit available to researchers but also lays the groundwork for future innovations aimed at realizing the full potential of direct reprogramming.