
The lack of standardised workflows and ambiguous metabolite annotations hampers metabolomics integration with prior knowledge, thus limiting the extraction of meaningful biological insights. We present MetaProViz (Metabolomics Processing, functional analysis and Visualization), an open-source Bioconductor R package for metabolomics data analysis that integrates prior knowledge to generate mechanistic hypotheses ( https://saezlab.github.io/MetaProViz/ ). MetaProViz operates on annotated intensity values and offers a flexible framework consisting of five modules: processing, differential analysis, prior knowledge integration, functional analysis and visualisation, applicable to intracellular and exometabolomics experiments. To improve functional analysis, we created the Metabolism Signature Database (MetSigDB), a collection of annotated metabolite sets. MetSigDB includes pathway-metabolite, metabolite-receptor, metabolite-transporter sets, and chemical class-metabolite sets. MetaProViz enables the conversion of gene sets to metabolite sets, metabolite identifier expansion and analyses mapping ambiguities. The MetaProViz functional analysis toolkit includes sample metadata analysis, enrichment analysis and biologically informed clustering. By applying MetaProViz to kidney cancer metabolomics data, we identified increased methionine usage in line with decreased methionine levels in tumour samples. In summary, MetaProViz facilitates and improves the analysis and interpretation of metabolomics data. MetaProViz is an open-source Bioconductor R package that integrates curated prior knowledge and metabolite annotation handling to enable reproducible metabolomics analysis, improve functional interpretation, and generation of mechanistic hypotheses from intracellular and extracellular metabolomics. MetaProViz is an open-source Bioconductor R package that integrates curated prior knowledge and metabolite annotation handling to enable reproducible metabolomics analysis, improve functional interpretation, and generation of mechanistic hypotheses from intracellular and extracellular metabolomics.
Transcription is an inherently dynamic and stochastic process that often occurs in bursts, governed by gene–gene regulatory interactions and thereby driving cell-to-cell heterogeneity. However, a genome-wide, mechanistic understanding of how regulatory networks globally shape transcriptional bursting dynamics remains lacking. Here, we present BurstLink, an interpretable and tractable statistical-mechanistic framework that simultaneously infers coupled regulatory interactions and transcriptional bursting kinetics at the genome-wide scale from single-cell data. BurstLink introduces reweighted mutual information to quantify regulatory strength as network edge weights, while jointly inferring regulatory directionality and interaction type for each gene pair within a unified mechanistic model of transcriptional bursting. Applied to mouse embryonic fibroblasts data, BurstLink reveals several genome-wide regulatory mechanisms on transcriptional bursting: downstream target genes exhibit higher burst frequency and gene-expression variability than upstream transcription factor genes; stronger transcription factor binding affinity is associated with lower burst frequency and higher burst size of target genes. Notably, positive regulation primarily enhances the burst frequency and gene-expression variability in target genes, in contrast to negative regulation. In summary, BurstLink deciphers multiple general principles of global transcriptional dynamics, providing novel biological insights into cell fate decisions. BurstLink jointly infers regulatory interactions and transcriptional bursting kinetics across an entire gene regulatory network from single-cell data, revealing genome-wide principles of how regulation shapes bursting. BurstLink jointly infers regulatory interactions and transcriptional bursting kinetics across an entire gene regulatory network from single-cell data, revealing genome-wide principles of how regulation shapes bursting.
The state of a cell depends not only on protein abundance, but also on the biochemical and cellular activities of proteins, which are largely invisible to abundance profiling alone. Here, we introduce a multi-omics framework that infers context-specific protein activities from transcriptomic, phosphoproteomic, and protein correlation-based protein-protein interaction data, integrating modality-specific algorithms via network diffusion. Applying it to a panel of phenotypically diverse HeLa cell lines, whose genetic drift provides a natural perturbation system, we make three findings. First, physical separation of monomeric and assembled protein fractions by protein correlation profiling provides direct evidence that complex assembly buffers variation in gene copy number and transcription, a mechanism previously only inferred from bulk measurements. Second, using Let7 perturbation data, CRISPR gene dependency scores, and subcellular localization, we orthogonally validate that inferred protein activities capture functional regulation linked to cellular phenotypes inaccessible from abundance data alone. Third, differential analysis of context-specific activity profiles identifies molecular mechanisms underlying phenotypic divergence, including a WIPF1/WIPF2–Arp2/3 axis governing invadopodium formation and infection susceptibility, and an immunoproteasome switch linked to immune adaptation. A multi-omics framework infers protein activities from co-fractionation, phosphoproteomic, and transcriptomic data. Applied to a panel of HeLa cell lines, it reveals how protein complex assembly buffers genetic variation and drives phenotypic divergence. A multi-omics framework infers protein activities from co-fractionation, phosphoproteomic, and transcriptomic data. Applied to a panel of HeLa cell lines, it reveals how protein complex assembly buffers genetic variation and drives phenotypic divergence.
Noncanonical small RNAs, such as tRNA-derived (tsRNAs) and rRNA-derived (rsRNAs) fragments, are more abundant than microRNAs and arise from selective cleavage events rather than random degradation. While fragmentation of parental RNAs produces functionally diverse small RNAs, current analytical approaches are limited to abundance measures and cannot systematically quantify differential cleavage signals. Here, we present qMAP, a computational framework profiling differential fragmentation of parental RNAs from small RNA sequencing data. qMAP integrates two complementary models to identify condition-specific fragmentation patterns and includes a dedicated module to pinpoint the small RNA species driving these differences. Using qMAP, we uncover dynamic tRNA and rRNA fragmentation during mouse cell reprogramming, demonstrate the classification power of RNA fragmentation in human ulcerative colitis, develop and validate a blood-based RNA fragmentation signature of recurrent implantation failure, and identify aging-associated RNA fragmentation in sperm, which supports RNA fragmentation as a distinct regulatory dimension beyond expression/abundance information. qMAP enables systematic exploration of the regulatory “RNA fragmentome”, providing a foundational tool for both mechanistic discovery and translational applications of noncanonical small RNAs. qMAP is a computational framework quantifying differential fragmentation of parental RNAs, which reveals “RNA Fragmentome” as a distinct regulatory layer. Across cell development, diseases, and sperm aging, qMAP uncovers RNA fragmentation dynamics with mechanistic and translational relevance. qMAP is a computational framework quantifying differential fragmentation of parental RNAs, which reveals “RNA Fragmentome” as a distinct regulatory layer. Across cell development, diseases, and sperm aging, qMAP uncovers RNA fragmentation dynamics with mechanistic and translational relevance.
Abstract Many bacterial species form self-organized macroscale patterns through swarming. Despite its extensive genetic tractability, Escherichia coli remains underexplored for robust, applied control of swarming. Here we develop a set of E. coli strains that generate centimeter-scale swarming patterns to spatially record environmental inputs. Specifically, we modulate the expression of swarming-related genes in response to chemical and optical signals, reshaping baseline swarm patterns in analog or binary-like fashions. To decode bacterial patterns across space and time, we develop scalable computational methods incorporating feature extraction, regression, and deep-learning models. Time-lapse imaging reveals that colonies record inputs dynamically, enabling early-stage classification. This work establishes a strategy for spatial information recording in E. coli and expands the toolkit for programming emergent microbial behaviors at macroscopic scales.
Metabolic modeling with stoichiometric models and flux balance analysis (FBA) has greatly advanced our understanding of metabolism. However, valid FBA predictions require mechanistically correct constraints. Thermodynamic constraints can increase the mechanistic foundations of stoichiometric models and reduce the solution space, but incorporating them has so far required cumbersome manual effort. To circumvent manual curation, we introduce 'Thermo-Flux', a semi-automated Python package that converts stoichiometric models into comprehensive thermodynamic-stoichiometric models. 'Thermo-Flux' enables (i) automated mass and charge balancing while considering physical and biochemical parameters, (ii) definition of transporter variants and Gibbs energies for transport processes, (iii) handling of metabolites with unknown structures or Gibbs energies, and (iv) integration of recent methods for determining Gibbs energies and their uncertainties. To guide users, we provide detailed instructions on how to use 'Thermo-Flux' and include background information to facilitate appropriate modeling assumptions. We highlight the applicability of 'Thermo-Flux' by converting 87 stoichiometric models from the BiGG database and demonstrate improved flux predictions for a genome-scale yeast model (iMM904). We expect 'Thermo-Flux' to support fundamental and applied metabolic research.
ATP-competitive kinase inhibitors represent one of the largest classes of targeted anti-cancer drugs. While their primary mechanism is to block catalytic activity, they can also trigger paradoxical phenotypic effects that cannot be explained by catalytic inhibition alone. These observations point to a hidden layer of drug action that modulates non-catalytic kinase functions via changes in kinase conformation and protein-protein interactions (PPIs). Here, we developed a multimodal proteomics approach combining limited proteolysis coupled mass spectrometry on affinity-purified samples (AP-LiP-MS), AP-MS, and proximity labeling-MS to map inhibitor-induced conformation and PPI changes. We show that inhibitor binding causes structural rearrangements in the autoinhibitory domains (AIDs) of all tested kinases, consistent with a transition to an open, active-like kinase conformation. These structural shifts drive distinct kinase-protein interaction changes that control non-catalytic functions: sequestration of AMPK by inhibited CAMKK2 blocks phosphorylation by other kinases, CHEK1 inhibition causes dissociation from the mitochondrial protein CLPB and leads to mitochondrial fragmentation, and structural changes in inhibited PRKCA trigger rapid relocalization to cell junctions. Thus, we identify the ATP-binding site as a major organizing center of kinase conformation and interaction. Our work suggests that these on-target, off-mechanism effects are likely to occur in other kinases as well, and provides the analytical framework to systematically characterize a frequently overlooked phenomenon highly relevant for understanding drug side effects to guide the development of novel therapeutics.
Although in the wild bacteria likely spend most of their time deprived of nutrients and slowly starving to death, very little is known about how bacteria adapt their phenotype to starvation. Here we combine microfluidics with quantitative fluorescence microscopy of transcriptional reporters to comprehensively quantify growth and gene expression at the single-cell level in E. coli during carbon starvation. We find that all cells immediately stop growing upon loss of carbon source and that almost all remain alive for over 30 hours. Furthermore, entry into starvation triggers a dynamic expression program that is remarkably homogeneous across single cells, but highly variable across genes, causing dramatic remodeling of the proteome early in starvation. We further show that, as protein production and the rate of protein degradation both decay approximately exponentially, protein concentrations become essentially 'frozen' after the first 5 to 10 hours, setting phenotypes for several days of starvation. Finally, using experiments in which gene expression is inhibited for different periods, we show that protein production in the first 5 hours is crucial for protecting cells from stresses late in starvation.
The gut microbiota is implicated in adverse effects associated with low-calorie sweeteners. Yet, the direct impact of sweeteners on gut bacteria remains largely uncharacterized. Here, we report interactions between 25 phylogenetically diverse gut bacterial strains and 39 commercially used sweeteners. We tested these sweeteners individually and in combination with four commonly co-consumed compounds, viz., advantame, caffeine, vanillin, and duloxetine. Three-quarters of the tested sweeteners individually impacted the growth of at least one tested bacterial strain. Further, over 100 interactions were found between sweeteners and the four co-consumed compounds. Isosteviol, a commonly used sweetener-component, and duloxetine, an antidepressant, synergistically inhibited Roseburia intestinalis, a bacterium previously linked to glucose homeostasis, and Parabacteroides merdae, a prevalent commensal linked to healthy microbiota. Proteomic, metabolomic, and genetic analyses indicate altered small molecule transport underpinning this sweetener-drug synergy. The isosteviol-duloxetine combination also modulated metabolism of a synthetic gut bacterial community, leading to increased toxicity to HeLa cells and altered secretion of inflammation-modulatory cytokines IL-6 and IL-8 by Caco-2 cells. Our data warrant further studies on interactions between low-calorie sweeteners and common xenobiotics.
Abstract Cross-linking mass spectrometry is a powerful method for structural analysis, but choosing between cleavable and non-cleavable cross-linkers remains challenging. We rigorously compared non-cleavable DSS with cleavable DSSO and found that DSS consistently yields more cross-link identifications from isolated protein complexes to bacterial lysates. The advantage of DSS diminishes as sample complexity increases. At the highest complexity tested—human cell lysate—the trend reverses, with DSSO outperforming DSS. The superior performance of DSS in less complex samples is likely explained by its longer and more flexible spacer arm, which interrogates a spatial volume >40% larger than that of DSSO. For both cross-linkers, the number of identified cross-links decreases as the search space expands, but more steeply for DSS. This sharper decline arises from DSS cross-links producing slightly lower fragment ion coverage, not from the absence of signature ions that could reduce search space. Fragment ion coverage is key to interactome mapping: when coverage reaches 85% or above, identification sensitivity hardly decreases as the search space expands, regardless of the cross-linker used. In summary, we recommend DSS for samples no more complex than bacterial lysates. For interactome mapping of mammalian cells, although DSSO outperforms DSS, neither achieves deep interactome coverage.
Understanding how gene regulatory networks (GRNs) dynamically orchestrate cell fate emergence remains a fundamental challenge. Here, we present GRNvelo, a computational framework that reconstructs multiscale cell fate dynamics by integrating GRNs with phenotypic dynamics from temporal single-cell RNA-seq data. GRNvelo establishes a biologically interpretable and mathematically rigorous multiscale model that couples GRN-driven single-cell velocity with nonlocal cell growth-mediated population dynamics. To operationalize this model, GRNvelo devises a two-phase cooperative optimization algorithm based on physics-informed neural networks (PINNs): TC-PINN for jointly inferring GRN velocity and latent time, and MP-PINN for refining GRN velocity within the context of cell population dynamics. In benchmark evaluations, GRNvelo demonstrates superior performance across two synthetic datasets and four real datasets, including branching development and diverse perturbation-response scenarios. Collectively, GRNvelo not only accurately infers GRN-driven cell fate dynamics but also predicts altered cell fates in response to diverse genetic perturbations, including dynamic and combined ones, thus establishing a new computational paradigm for predicting and modulating cell fate outcomes.
Cold tumors like pancreatic cancer suffer from poor immune infiltration, limiting effective anti-tumor responses. The chemokine CXCL9 promotes immune cell recruitment, but the signaling mechanisms regulating its expression in tumor cells remain poorly understood and underexplored as targets for modulation. We present a framework that integrates active learning with mechanistic logic-ODE models to guide perturbation screenings and uncover regulators of CXCL9 in pancreatic cancer cells. Using perturbation-response data and curated prior knowledge, we trained interpretable models to identify signaling mechanisms that enhance CXCL9 expression and prioritize drug combinations. Active learning enabled data-efficient model refinement and guided informative experiments under resource constraints. Benchmarking on synthetic data and experimental validation confirmed the performance of different acquisition strategies and its applicability to feasible iterative wet lab experiments. Our results demonstrate how combining active learning with mechanistic modeling supports rational, targeted experimental design.
Abstract In controlled proof-of-concept studies, healthcare artificial intelligence (AI) systems now routinely match or exceed specialist-level performance. Despite these impressive technical achievements, most systems remain confined to research settings, highlighting an opportunity to better prepare academic AI research for subsequent clinical translation. Successful translation requires attention to technical, regulatory, organizational, and infrastructural factors, ideally from the earliest stages of research. This Perspective focuses on trustworthiness as a key factor that academic researchers can address proactively. First, we outline how trustworthiness requirements differ across three stakeholder groups—patients, clinicians, and regulatory bodies—and how academic research can better align with these expectations. We then discuss key dimensions of trustworthy AI—explainability, uncertainty quantification, and evaluation—and how researchers can address them during study design. We also consider foundation models, which require adapted validation approaches. Based on this analysis, we propose a stakeholder-centered framework that emphasizes proactive integration of trust requirements during early development phases rather than post hoc considerations.
Advances in mass spectrometry have transformed plasma proteomics, allowing high-throughput analysis of large cohorts. This study utilized the TIMES (Tracking Individuals Monthly for Evaluating Stability) cohort, consisting of 51 healthy participants monitored monthly over 12 months, to evaluate intra- and inter-individual variability in the plasma proteome. Approximately 600 samples were analyzed in two laboratories independently, revealing strong correlations despite methodological differences. The study revealed high stability of the plasma proteome within donors over a year, with larger differences observed between donors. A support vector machine model achieved a 98% median classification accuracy, confirming stable, donor-specific proteomic profiles. Immunoglobulins, often overlooked in plasma proteomics, were found to contribute significantly to donor-specific signatures, showing pronounced inter-individual variability but remarkable intra-donor stability. In contrast, C-reactive protein and other inflammation markers exhibited significant temporal fluctuations from donor-specific baseline levels. This inter-laboratory study highlights the importance of longitudinal sampling for biomarker discovery, the robustness of MS-based proteomic workflows, and provides new insights into immunoglobulin dynamics and variability in human plasma.
Complex genome rearrangements are intriguing because it remains unclear whether they occurred gradually or suddenly, and what promoted or impeded their development. We started by asking whether double loss of heterozygosity (dLOH) involving two different chromosomes is simply the result of two independent single LOH events. We examined hundreds of vegetatively growing diploid yeast strains with gene deletions known to stimulate LOH through different mechanisms. In dozens of these strains, dLOH mutants were overrepresented by tens or even hundreds of times. Interestingly, the (deleted) genes that caused the greatest instability were functionally diverse yet apparently affected the same biological process: cell cycle control. Furthermore, mutants were dominated by aneuploidy, which often involved multiple non-target chromosomes, with virtually no segmental loss, non-homologous recombination, or other small-scale changes. We conclude that the supply of separate, local rearrangements affecting only the required regions is insufficient, even in mutator strains. Consequently, double-pointed selection tends to reveal extensive reassortments resulting from systemic failures in mitotic cell division. Therefore, overcoming multiple adaptation barriers tends to produce aneuploidy, which is often reversible.
Understanding the mechanism of action (MoA) of bioactive compounds is a central challenge in drug discovery and chemical biology. We propose a strategy for integrating morphological data with proteomics to provide deeper insights into the MoA of compounds. We combine the rich phenotypic profiles of Cell Painting (CP) with the unbiased protein target detection of Thermal Proteome Profiling (TPP), and construct protein-protein interaction networks based on potential targets identified using both assays. We validated our method with public TPP datasets for the five compounds (+)-JQ1, I-BET151, Vemurafenib, Crizotinib and Panobinostat, and public Proteome Integral Solubility Alteration (PISA) data for 49 compounds, together with CP data for 5259 drugs on U2OS cells. We show that the combined approach could accurately identify known MoAs for four out of five validation compounds and known targets for 29 out of 49 compounds. Finally, we deployed our method to characterize sinomenine, a compound with elusive knowledge of MoA. Our findings revealed novel facets of sinomenine's biological activity and highlight the value of multimodal profiling for chemical biology.
Understanding how cells transition between states is central to biology. RNA velocity has emerged as a transformative framework for inferring future cellular states from unspliced and spliced mRNA transcripts. A new deep generative and agent‑based method can decode and manipulate cellular dynamics within native tissue by dissecting spatially resolved RNA velocity and simulating regulatory interventions, opening new opportunities to explore how tissue context shapes and potentially controls cell‑state transitions. This News Views highlights a recent study by Raghavan and colleagues (in this issue of MSB) that introduces veloAgent, a computational framework that integrates transcriptional dynamics with spatial information to dissect spatially resolved RNA velocity.
Traditional methods for engineering and sequence-fitness analysis of proteins in mammalian cells are limited by the time, cost, and labor associated with plasmid cloning and preparation. Here we present Microbe-Independent Deep Assembly and Screening (MIDAS), a deterministic platform for rapid protein variant expression and characterization in mammalian cells that bypasses microbial cloning by directly transfecting PCR-assembled genes. MIDAS enables high-quality sequence-fitness assessment of arbitrary mutational spaces, including truly deep saturation mutagenesis and combinatorial variant assembly, requiring less than one workday from initial PCR to cell transfection. Using MIDAS, we engineer a high-performance acetylcholine neurotransmitter bioluminescent indicator (ACh-NeuBI), achieving stepwise improvements in responsivity through linker engineering, single-site, and multi-site mutagenesis. We also apply MIDAS to engineer improved NanoLuc luciferase variants for multiple substrates, and to characterize the structural basis of mutational tolerance and substrate specificity. Thus, MIDAS is a versatile method for rapid plasmid-free protein engineering and sequence-fitness analysis in mammalian cells, offering a practical alternative to cloning-based approaches for many protein optimization and characterization tasks.
Bulk high-resolution mass spectrometry provides sensitive and global snapshots of metabolites involved in cancer metabolism. However, intratumoral heterogeneity obscures the cellular origins of detected metabolites, making it difficult to identify reproducible and predictive metabolic markers. Here, we present "Spatially guided MEtabolomics (SgME) profiling", a multi-modal metabolomics data analysis approach that delineates and maps metabolic regions (MERs), including overlapping regions, within tumor tissues. We applied SgME profiling to human hepatocellular carcinoma (HCC) tumors and refined potential RNA markers that were also found in previous transcriptomics or bioinformatics studies to those specifically associated with malignant regions. We further estimated that more than 50% of the highly abundant metabolites detected in bulk tumors originated from the non-malignant MERs and are therefore unlikely to be predictive and/or reproducible markers. Importantly, SgME profiling also revealed new potential metabolic markers that were not apparent in bulk analysis because they increased in low-grade tumor regions but declined sharply in necrotic regions. Together, these findings show that SgME profiling overcomes key limitations of conventional metabolomics profiling by enabling more granular, spatially resolved metabolic characterization.