Abstract Mechanistic and functional analysis of omics data largely relies on the incorporation of prior knowledge; however, connecting metabolomics data and knowledge is a major methodological challenge. This is largely driven by the diverse prior knowledge being fragmented across many databases requiring the merging of different database records across chemical structures, identifiers, and varying levels of structural specificity. Hence, this limits mechanistic interpretation and functional characterisation of the metabolome. Here, we present OmniPath Metabo, a comprehensive, harmonized, metabolome-centric database covering metabolites, lipids, food-derived compounds, and small molecule drugs, along with their associated receptors, transporters, enzymes, reactions, allosteric regulators, and disease associations. OmniPath Metabo harmonizes attributes using controlled vocabularies and ontologies, structures and built-in cheminformatics to map identifiers and track ambiguity. OmniPath Metabo is built directly from 40+ original resources and is freely accessible via an interactive web app and API at metabo.omnipathdb.org . OmniPath Metabo enables dynamic, context-specific construction of subnetworks to serve dedicated purposes, such as cell-cell communication or integrated multi-omics metabolite-driven regulation, connecting reactions, allosteric regulation, metabolite-receptor and metabolite-transporter interactions. Combining it with the over 170 other resources in OmniPath, it can be used for integrated networks of signaling, gene regulation, and metabolism. We showcase the application of OmniPath Metabo by analysing publicly available metabolomics data of lung cancer cell lines and metabolic footprints to mutational patterns. In summary, OmniPath Metabo transforms fragmented resources into a harmonised prior knowledge framework for a mechanistic and functional analysis of the metabolome.
The web is increasingly accessed by AI agents rather than humans. Every agent needs knowledge, especially in the life-sciences, where agentic pipelines are growing fast. Access to the literature is a crucial part of that need, and resources such as Europe PMC, with over 40M indexed records, are widely used to meet it. Yet these resources were not built for AI agents: they take keywords and complex syntax and return whole papers, so every agent must learn the syntax, issue several searches, and read full papers to find the evidence it needs. We introduce EMBL AI Librarian, a knowledge layer that upgrades the Europe PMC interface for AI agents: an agent asks in natural language and receives evidence that answers it. A single LLM orchestrates the whole knowledge retrieval process: it plans complementary subqueries executed by the live Europe PMC search engine, then reads the selected papers and locates the relevant evidence. We evaluate Librarian across four benchmarks: literature synthesis, claim verification, open-domain question answering, and downstream biology tasks such as protocol questions and sequence manipulation. On ScholarQABench, Librarian improves Citation F1 by more than 16 points over strong recently published baselines. Used as the retrieval layer of an existing claim-verification pipeline, it increases agreement with expert consensus; and on the open-form LitQA2 benchmark, a GPT-5.4 agent scores about 8 points higher when grounded in Librarian than with web search. Overall, our results show that equipping life-science agents with the Librarian knowledge layer improves performance across a range of tasks. We release our code publicly at https://github.com/petroni-lab/librarian
Abstract Biomedical discovery is hindered by fragmented, modality-specific repositories and uneven metadata, limiting integrative analysis, accessibility, and reproducibility. To address these challenges, we present CROssBARv2, a provenance-rich biomedical data-and-knowledge integration platform that unifies heterogeneous sources into a maintainable, scalable system. By consolidating diverse data types into an extensive knowledge graph enriched with standardised ontologies, rich metadata, and deep learning–based vector embeddings, CROssBARv2 alleviates the need for researchers to navigate multiple siloed databases and can facilitate downstream tasks, including predictive modelling and mechanistic reasoning, enabling applications such as drug repurposing and protein function prediction. The platform offers interactive graph exploration and embedding-based semantic search with CROssBAR-LLM, an intuitive natural language question-answering system that grounds large language model (LLM) outputs in the underlying knowledge graph to mitigate hallucinations. We assess CROssBARv2 through (i) multiple use-case analyses to test biological coherence and relational validity; (ii) knowledge-augmented biomedical question-answering benchmarks comparing CROssBAR-LLM against generalist LLMs; and (iii) a deep learning–based predictive modelling experiment for protein function prediction leveraging the heterogeneous structure of CROssBARv2. Collectively, CROssBARv2 provides a scalable, AI-ready, and user-friendly foundation that facilitates hypothesis generation, knowledge discovery, and translational research.
Deletion of chromosome 5q [del(5q)] is the most common cytogenetic abnormality in myelodysplastic neoplasms (MDS) and results in haploinsufficiency of multiple genes, including CSNK1A1. Recurrent CSNK1A1 mutations, predominantly affecting the E98 hotspot, occur almost exclusively in del(5q) MDS and are associated with adverse outcomes, yet their impact on CK1ɑ function remains unclear. Using integrated transcriptomic, (phospho)proteomic, and kinome activity profiling in hematopoietic stem and progenitor cells (HSPCs), combined with in vivo serial transplantation assays, we show that Csnk1a1 E98V represents a change-of-function rather than a loss-of-function mutation. Unlike Csnk1a1 haploinsufficiency, Csnk1a1 E98V preserves long-term hematopoietic reconstitution and does not enhance clonal expansion in vivo. Instead, the mutation induces suppression of kinase signaling networks, leading to coordinated repression of ribosomal gene expression, protein translation, and cell cycle programs. This signaling rewiring is accompanied by metabolic reprogramming characterized by reduced mitochondrial respiration, increased glycolytic flux, and an inability to adapt to metabolic challenges, creating a stress-tolerant but inflexible cellular state. Notably, Csnk1a1 E98V cells exhibit impaired megakaryopoiesis and increased vulnerability to iron overload, as well as RSL-3-mediated ferroptosis. Analysis of del(5q) MDS patients confirmed that CSNK1A1 mutations are associated with distinct clinical features, including thrombocytopenia, elevated myeloblasts, and reduced bone marrow iron levels. Together, our findings support a two-step model in which del(5q)-associated CSNK1A1 haploinsufficiency drives clonal expansion, followed by acquisition of CSNK1A1 mutations that promote stress tolerance rather than increased proliferation. This adaptive rewiring exposes metabolic and iron-dependent vulnerabilities that may be therapeutically exploited.
Language models and agents are increasingly used in biomedicine, but current benchmarks reward correct answers even when the underlying reasoning is flawed. Here we introduce Karenina, an open-source framework that turns expert knowledge into multi-dimensional evaluations of questions, conversations and autonomous agents. Illustrated in Question-Answer pairs, multi-turn conversations and autonomous data-analysis, these dimensions together moves evaluation beyond scoring, enabling trustworthy decision-making with AI in biomedicine.
Precision oncology aims to tailor treatment according to tumor-specific molecular alterations, but the success of aberration-guided therapies has been limited in clinical trials. Here, we develop an integrated whole-genome and transcriptome workflow to systematically distinguish functionally credible, predictive driver aberrations from non-functional alterations across all classes of genomic events. We applied the integrated omics workflow to 335 patients with ovarian high-grade serous carcinoma (HGSC) enrolled in the observational DECIDER study. Tumor samples were collected from multiple cancer sites as part of the standard cancer care. DNA and RNA were extracted together from snap-frozen tumor samples and sent to whole-genome and transcriptome sequencing. Sequencing data were processed with the Anduril 2 pipeline for detection and validation of short somatic changes and with the HMW toolkit and the nf-core/rnafusion pipeline for assessment of structural changes. Aberration-specific drug sensitivity was tested in patient-derived organoids with a drug screen combining targeted agents and chemotherapy. Using an agnostic integrated omics analysis, we identified clinically relevant ESCAT Tier II–III alterations in more than 40
Abstract A fundamental design pattern in biomolecular studies is to assay the same set of samples (organisms, tissue biopsies, or individual cells) by multiple different ‘omics assays. Group Factor Analysis (GFA) and its adaptation to high-dimensional settings, Multi-Omics Factor Analysis (MOFA), are widely used as a first-line approach to analyze such data and are effective in detecting patterns of correlation, organize them into so-called latent factors, and identify common and assay-specific factors. However, in many applications, a subset of the found factors just rediscovers already known covariates (e.g., disease subtypes, environmental covariates) while others may represent genuine novelty. Here, we present Semi-supervised Omics Factor Analysis (SOFA), a method that incorporates known covariates into the model upfront and focuses the factor discovery on novel sources of variation. We show SOFA’s effectiveness for discovering novel patterns by applying it to cancer, brain development and heart failure multi-omic data sets.
Recent advances in spatial omics technologies have provided unprecedented insight into tissue spatial organization, but challenges remain in aligning spatial slices and integrating complementary single-cell and spatial data. Here, we propose TOAST (topography-aware optimal alignment of spatially resolved tissues), an optimal transport (OT)-based framework that extends the classical fused Gromov-Wasserstein (FGW) objective to more comprehensively model the heterogeneity of local molecular interactions. By introducing “spatial coherence,” quantified through the entropy of local neighborhoods, and “neighborhood consistency,” which preserves the expression profiles of neighboring spots, TOAST’s objective function improves the alignment of spatially resolved tissue slices and the mapping between single-cell and spatial data. Through comprehensive evaluations, we demonstrate that our method consistently outperforms traditional FGW and other OT-based alignment methods. By integrating spatial constraints into OT, our framework provides a principled approach to enhance the biological interpretability of spatially resolved omics data and facilitate multimodal data integration.
The lack of standardised workflows and ambiguous metabolite annotations hampers metabolomics integration with prior knowledge, thus limiting the extraction of meaningful biological insights. We present MetaProViz (Metabolomics Processing, functional analysis and Visualization), an open-source Bioconductor R package for metabolomics data analysis that integrates prior knowledge to generate mechanistic hypotheses ( https://saezlab.github.io/MetaProViz/ ). MetaProViz operates on annotated intensity values and offers a flexible framework consisting of five modules: processing, differential analysis, prior knowledge integration, functional analysis and visualisation, applicable to intracellular and exometabolomics experiments. To improve functional analysis, we created the Metabolism Signature Database (MetSigDB), a collection of annotated metabolite sets. MetSigDB includes pathway-metabolite, metabolite-receptor, metabolite-transporter sets, and chemical class-metabolite sets. MetaProViz enables the conversion of gene sets to metabolite sets, metabolite identifier expansion and analyses mapping ambiguities. The MetaProViz functional analysis toolkit includes sample metadata analysis, enrichment analysis and biologically informed clustering. By applying MetaProViz to kidney cancer metabolomics data, we identified increased methionine usage in line with decreased methionine levels in tumour samples. In summary, MetaProViz facilitates and improves the analysis and interpretation of metabolomics data. MetaProViz is an open-source Bioconductor R package that integrates curated prior knowledge and metabolite annotation handling to enable reproducible metabolomics analysis, improve functional interpretation, and generation of mechanistic hypotheses from intracellular and extracellular metabolomics. MetaProViz is an open-source Bioconductor R package that integrates curated prior knowledge and metabolite annotation handling to enable reproducible metabolomics analysis, improve functional interpretation, and generation of mechanistic hypotheses from intracellular and extracellular metabolomics.
Lipid transfer proteins (LTPs) maintain the specialized lipid compositions of organellar membranes1,2. In humans, many LTPs are implicated in diseases3, but the cargo and auxiliary lipids that facilitate the transfer of the majority of LTPs remain unknown. Here we combined biochemical, lipidomic and computational methods to systematically characterize LTP-lipid complexes4 and measure how LTP gains of function affect cellular lipidomes. We identified bound lipids for around half of the hundreds of LTPs that we analysed, confirming known ligands and identifying new ones across most LTP families. Gains in LTP function affected the cellular abundance of both their known and newly identified lipid ligands, indicating comparable functional relevance of the two ligand sets. Using structural bioinformatics, we characterized mechanisms that contribute to lipid selectivity and identified preferences based on headgroup or acyl chain. We demonstrate some basic principles of how LTPs mobilize their ligands. They commonly interact with several classes of lipids and exhibit broad but selective preference for particular headgroups and for lipid species with shorter acyl chains that contain one or two unsaturated carbons, suggesting that only subsets of lipid species are efficiently mobilized. The datasets represent a resource for further analysis in different cell types and states, such as those associated with pathologies.
Systemic lupus erythematosus (SLE) shows marked clinical and molecular heterogeneity, yet patient stratification often relies on gene expression signatures lacking multicellular context. Here we construct a transcriptional patient map of SLE by analyzing 1,167 total samples (783 SLE, 384 healthy controls) across different resolutions, including single-cell and bulk blood as well as spatially resolved kidney tissue transcriptomes. Using an unsupervised approach we inferred patient-level transcriptomic immune programs from two independent single-cell RNA sequencing cohorts of peripheral blood mononuclear cells (PBMCs), capturing both differences between SLE and health as well as within-SLE heterogeneity. Specifically, we identified four conserved programs comprising two multicellular inflammatory programs driven by interferon and TNF/NFkB activity across immune cells, and two cell type-specific programs reflecting CD8 T cell cytotoxicity and a CD4 T cell naive-to-effector state. Functional analysis of these programs revealed a rewiring of both cell-to-cell interactions and task allocation across cell types during disease activation. In addition, mapping these programs onto an external longitudinal blood transcriptomic cohort predicted flare risk and identified candidate blood protein biomarkers detectable by proteomics. Finally, we showed that these blood programs were enriched in immune-infiltrated glomerular regions from kidney biopsies of individuals with lupus nephritis using spatially resolved transcriptomic data, thereby linking systemic immune programs to local tissue pathology.
Tomato (Solanum lycopersicum), despite being the most important vegetable crop world-wide, remains vulnerable to over 200 diseases caused by different pests. Although tomato molecular response to individual stresses is well studied, the gene regulatory network (GRN) representing the crosstalk and trade-offs of multi-stress responses remains almost unexplored. We developed GENIAL (Gene rEgulatory Network and topologIcal datA anaLysis) to refine and analyze complex GRNs and TomTom, a knowledge graph that gathers 11 publicly available databases in a unique FAIR (Findable, Accessible, Interoperable, Reproducible) resource. To test GENIAL, we used transcriptomics data from tomato subjected to six distinct pathogens from the literature. GENIAL yielded the identification of transcription factors (TFs) coordinating the specific and multiple pathogen response. Functional validation using the virus-induced gene silencing system in tomato demonstrated that ETHYLENE RESPONSE FACTOR 16 and TCP DOMAIN PROTEIN 17 act as key TFs in the response to Botrytis cinerea, as silencing of either TF resulted in increased susceptibility. The validation of selected downstream targets allowed the validation of the robustness of the interactions highlighted by GENIAL. This study represents a proof of concept of our framework and can be extended to include other molecular layers and scaled to other questions involving tomato and beyond.
Metastasis remains the leading cause of cancer-related mortality and is driven by pronounced tumour cell plasticity1. Here we identify the transmembrane glycoprotein trophoblast cell-surface antigen 2 (TROP2) as a marker of poor-prognosis colorectal cancer (CRC) associated with WNTlow, fetal-like tumour cell states that are linked to metastasis and therapy resistance. Functional analyses demonstrate that TROP2+ cells exhibit context-dependent stem-like capacity and the ability to initiate metastatic outgrowth. Given that these detrimental tumour states converge on the cell-surface antigen TROP2, we explored therapeutic targeting of this cell population using clinically relevant TROP2-directed antibody-drug conjugates. Time-resolved analyses reveal therapy-associated dynamics in tumour cell state composition between WNThi LGR5+ states and WNTlowTROP2+ fetal-like states. Conventional chemotherapy promotes the induction of TROP2-expressing cells, whereas TROP2 antibody-drug conjugates selectively target these populations and remodel the tumour cell state landscape. Exploiting this plasticity, combined chemotherapy and TROP2 targeting enhances anti-tumour efficacy in patient-derived models. Together, our findings identify TROP2 as a therapeutic vulnerability of CRC and highlight the importance of targeting tumour cell states to improve therapeutic efficacy and overcome resistance in advanced disease.
Medulloblastoma, the most common malignant brain tumor of childhood, exhibits significant biological complexity that demands deeper exploration. Here, we present a large multiomics dataset integrating data from 384 primary medulloblastoma patient samples across five omic layers: CpG methylome, transcriptome, proteome, phosphoproteome, and metabolome, paired with associated clinical metadata. Data integration revealed intertumoral heterogeneity of lipid metabolism across proteomic subtypes. Notably, while the MYC-FASN-SCD axis drives lipid biosynthesis, pathway inhibition elicits a compensatory escape mechanism in vivo through exogenous fatty acid uptake. Unexpectedly, we demonstrated that MYC triggers lipid storage, creating a unique dependency on lipid droplet-mitochondria communications to sustain tumor maintenance in vivo. Together, this comprehensive analysis reveals a targetable vulnerability downstream of MYC that constitutes a promising therapeutic approach to treat currently untreatable medulloblastoma subtypes.
Cure rates for childhood malignancies using established therapy protocols have increased to an average of 80% but have reached a plateau. Moreover, survival rates are particularly low for some pediatric tumors—such as high-risk group 3 medulloblastomas, osteosarcomas, Ewing sarcomas, high-risk neuroblastomas, and high-grade gliomas—and dismal for patients with relapsed malignancies. A functional drug response profiling platform for pediatric solid and brain tumors has been established within the INFORM program to identify patient-specific vulnerabilities and biomarkers and to unravel molecular mechanisms associated with drug response profiles for clinical translation. In this study, we performed a multiomics analysis using drug sensitivity profiles, as well as genomic and transcriptomic data, of 81 pediatric solid tumor samples. The integrative analysis suggested two multiomics signatures associated with drug sensitivity. One signature distinguished neuroblastoma samples with sensitivity to navitoclax, a BCL2 family inhibitor. A second signature was specific to a subset of Wilms tumors harboring the SIX1 (Q177R) hotspot mutation that displayed high expression of MGAM, PTPN14, STAT4, and KDM2B and high sensitivity to MEK inhibitors. A patient-specific causal interaction network analysis suggested possible molecular interactions between MEK inhibitors and the SIX1 mutation in Wilms tumor samples. In conclusion, the integration of drug sensitivity profiling and multiomics data revealed potential biomarkers that may be associated with drug sensitivity in pediatric solid tumors. Patient-specific causal interaction network analysis further elucidated the interaction between inhibitors and signature biomarkers, providing insights that may inform clinical translation. Significance: The combination of multiomics analysis and drug sensitivity profiling identified two signatures related to drug sensitivity in pediatric solid tumors, contributing to the advancement of functional precision medicine and personalized treatment strategies. This article is part of a special series: Driving Cancer Discoveries with Computational Research, Data Science, and Machine Learning/AI .
Abstract Computational modeling provides a powerful framework for in silico exploration of anti-cancer therapeutic targets and tumor response mechanisms. Oncogenic signaling pathways play a central role in tumor behavior and represent promising targets for personalized combination therapies. However, these pathways are complex, and although logic-based models are well suited for representing signaling dynamics, they are often constrained by model-specific data requirements, limited scalability, and time-consuming manual curation. Here, we introduce Functional Integration of Contextualized Omics for Unraveling regulatory dynamicS (FICUS), a framework that integrates omics-driven network contextualization with dynamic Boolean and logic-ODE modeling. FICUS enables automated, data-driven protein network inference and patient stratification, allowing shared signaling mechanisms to be identified across patient subgroups while preserving patient-specific dynamic responses. We applied FICUS to the SU2C-MARK lung cancer cohort and the The Cancer Genome Atlas kidney cancer cohort, demonstrating its utility for post-hoc analyses and downstream interrogation of dynamic tumor models. Overall, our results highlight the flexibility of FICUS in capturing heterogeneous signaling mechanisms across patient subgroups, addressing a key challenge in precision oncology.
Signaling pathways are useful models for interpreting molecular data, but their coverage has long been constrained by classic biochemistry methods. The growing corpus of kinase-substrate interactions, coupled to phosphoproteomics improvements, pave the way to revisit classic signaling pathways. In this study, we explore context-specific signaling pathway inference from phosphoproteomics and kinase-substrate networks. Focusing on epidermal growth factor (EGF), we conduct a meta-analysis and generate three datasets representing the most comprehensive characterization of the EGF response to date. We infer kinase-kinase pathways and compare them to different ground truth sets. Literature-curated networks consistently yield the highest recovery of ground-truth interactions, with modest gains from network propagation methods. Up to 90% of interactions are absent from current ground truth sets, indicating many unexplored interactions supported by data and knowledge. Our results demonstrate the limitations of traditional views on signaling pathways and point to opportunities for generating better mechanistic hypotheses.
MOTIVATION:Chromatin 3D folding creates numerous DNA interactions, participating in gene expression regulation. Single-cell chromatin-accessibility assays now profile hundreds of thousands of cells, challenging existing methods for mapping cis-regulatory interactions. RESULTS:We present CIRCE, a fast and scalable Python package to predict cis-regulatory DNA interactions from single-cell chromatin accessibility data. CIRCE re-implements the Cicero workflow to analyse single-cell atlases, cutting runtime and memory use by several orders of magnitude. We also provide new options to compute metacells, grouping similar cells to reduce data sparsity. We benchmarked CIRCE against Cicero on two datasets of different sizes and demonstrated the improvement from CIRCE's metacells' strategy with promoter capture Hi-C data. We also evaluated how DNA interaction predictions are impacted by different pre-processing. We observed a negative impact of Cicero's count normalization, and the best performance was obtained with the single-cell count matrix directly. Finally, we demonstrated the scalability of CIRCE by processing a dataset of more than 700 000 cells and 1 million DNA regions in less than an hour. CIRCE should greatly facilitate the prediction of DNA region interactions for scverse and Python users, while providing new and up-to-date pre-processing insights. AVAILABILITY AND IMPLEMENTATION:CIRCE is released as an open-source software under the AGPL-3.0 licence. The package source code is available on GitHub at https://github.com/cantinilab/CIRCE, and its documentation is accessible at https://circe.readthedocs.io. The code to reproduce the presented results is available as a Snakemake pipeline at https://github.com/cantinilab/circe_reproducibility.s.
A transparent evaluation and proof of reproducibility, generalization and replicability of algorithms are the bedrock of method development in computational biology. Many benchmarking efforts have been developed for problems ranging from structural biology to translational biomedicine. Rigor is relatively controllable for tasks such as the prediction of patient outcomes or the outcomes of biological assays, but the problem is exacerbated when the aim is to benchmark foundation models. The parameters constituting them are supposed to capture the patterns underlying the data; therefore, the models are parameterized embodiments of the phenomena that gave rise to the data. How can we test the limitations of these models? Here, we discuss the epistemological value of foundation models; whether they can be refuted, verified or evaluated primarily on the basis of utility; what principles should guide their benchmarking; and what role the scientific community should play in that benchmarking process.