Tissue homeostasis and disease emerge from cell–cell interactions operating across spatial scales: from autocrine and juxtacrine signals within micrometers to paracrine gradients coordinating responses across tissues. While these can be read out from spatial transcriptomics, existing computational methods capture either local adjacency-based or long-range dependencies, but rarely both within a single framework. We introduce InterScale, a graph-transformer approach that jointly models local and global cellular interactions from spatial transcriptomics data. By integrating a Graph Convolutional Network as a local component with a global transformer encoder, InterScale learns multi-scale representations of cellular communication. A downstream workflow enables scale-resolved interpretation of interactions from gene to tissue level. Applied to Sonic Hedgehog morphogen patterning in neural organoids, InterScale resolves spatially restricted neuronal differentiation programs and broader progenitor regulatory states along the morphogen gradient. In a human pancreatic dataset contrasting healthy and type 1 diabetic tissue, it reveals disease-associated spatial reorganization and tissue remodeling. InterScale’s modular architecture supports diverse spatial transcriptomics platforms and provides a scalable, unbiased, and biologically interpretable framework for studying cellular interactions across scales.
Neurological disorders represent a growing global health burden requiring long-term, interdisciplinary rehabilitation. Computational neurorehabilitation (compNR) - the use of data-driven and model-based approaches to personalize treatment - offers new opportunities for precision rehabilitation. However, its clinical deployment is limited by fragmented data systems, poor interoperability, and low clinician engagement in model development. We embed the learning health system (LHS) framework in Neurorehabilitation through integration of multimodal data collection, model computation, and clinical visualization that enables clinician-ML collaboration in everyday neurorehabilitation practice. The system facilitates structured digital data capture, secure computational processing, and interoperable visualization of patient trajectories. Through a real-world deployment in stroke rehabilitation, we demonstrate how such an infrastructure bridges the gap between research models and clinical use, showcasing one approach to a translational pathway for compNR.
Single-cell sequencing technologies reveal cellular heterogeneity at high resolution, advancing our understanding of biological complexity. As datasets start to scale to tens of millions of cells, computational workflows face substantial bottlenecks, with CPU-based analytical pipelines requiring hours or days for routine processing steps like filtering, normalization, and clustering. These scalability limitations fundamentally restrict common interactive data exploration and iterative hypothesis testing. Here we introduce rapids-singlecell, a GPU-accelerated framework that integrates natively with the scverse ecosystem and operates directly on the AnnData data structure, which delivers orders-of-magnitude speedups for single-cell workflows. Built on CuPy arrays and the NVIDIA CUDA-X Data Science (RAPIDS) ecosystem, rapids-singlecell provides near drop-in GPU replacements for core scanpy-based analysis steps. Across standard single-cell workflows such as preprocessing, dimensionality reduction, neighborhood graph construction, clustering, and batch correction, rapids-singlecell achieves speedups of up to several hundred-fold compared to optimized CPU baselines. This reduces analysis time from hours to minutes on standard hardware, while maintaining consistent biological interpretations. These performance improvements make it possible to analyze large data sets in close to real time, without the need for data splitting. Together with real-time parameter tuning and iterative workflows, rapids-singlecell makes interactive large-scale single-cell analysis possible.
Large perturbation models require training data encompassing chemical, cellular, and assay diversity. Current transcriptomic resources for small-molecule modeling, however, are fragmented across technologies, metadata conventions, controls, doses, and preprocessing pipelines. We introduce Chem-PerturBridge, a harmonized multi-dataset resource comprising over 37k compounds, 136 cellular contexts, and 1.25M transcriptomic samples across eight assay types, with standardized identifiers, metadata, and replicate-aware condition-level effects. We use the resource to evaluate matched-condition agreement across datasets and replicate agreement within datasets. Matched same-compound conditions generally show weak agreement in fine-grained logFC rankings and magnitudes across most dataset pairs, often falling below same-context different-compound baselines. In contrast, logFC direction agreement is substantially more stable and usually exceeds these baselines. We further evaluate Chem-PerturBridge as a pretraining resource for compound representation learning. Under a compound-held-out OP3 evaluation split, embeddings pretrained on Chem-PerturBridge improve over L1000-only embeddings, Morgan fingerprints, and the descriptor-free OP3 baseline across metrics. An extensive molecule-holdout evaluation across 11 datasets further shows that models trained on Chem-PerturBridge outperform or match those that are not. Chem-PerturBridge therefore supports both diagnostic evaluation of cross-dataset signature agreement and model-oriented reuse of heterogeneous perturbation transcriptomic data.
Transcriptomics enables comprehensive, multiplexed characterization of cellular states, yet prevailing methods typically require cell fixation or lysis, precluding longitudinal analysis of RNA expression in living cells. Here, we present non-destructive transcriptomics by vesicular export (NTVE), a platform for multi-time-point monitoring of RNA expression dynamics in living cells. Stabilized RNA reporter barcodes can be selectively packaged and exported from cells via virus-like particles (VLPs) bearing bioorthogonal affinity handles for convenient multichannel tracking of co-cultured cells. Using an engineered poly(A)-binding protein adapter, NTVE exports endogenous transcripts from inducible human and murine cell lines with high concordance to conventional lysate-derived RNA-seq. NTVE captures transcriptome changes in response to genetic and chemical perturbations within the same cells over time using standard sequencing workflows. NTVE can further be equipped with fusogens to deliver mRNA-encoded effectors or ribonucleoprotein gene editors from sender cells, activating gene reporters in co-cultured recipient cells. We demonstrate the utility of NTVE for monitoring hiPSC differentiation through daily non-destructive transcriptomic profiling of lineage-specific marker dynamics.
B cell depleting therapies in multiple sclerosis (MS) have transformed disease management, yet the immunological mechanisms linking B cells to chronic compartmentalised CNS inflammation, the key pathological driver of disability accrual, remain poorly defined. Here, we leverage anti-CD20 therapy as an in vivo perturbational probe to disentangle MS immunobiology. We combine longitudinal, high-dimensional multimodal immune profiling in patients initiating ocrelizumab with a novel machine learning pipeline to resolve treatment-induced shifts in continuous immune cell states, revealing treatment-associated modulation of shared biological processes that act across the boundaries of discretely partitioned cell types. We identify a distinct chronically activated, proinflammatory-cytotoxic T cell state with CNS-homing properties that is selectively depleted following B cell ablation. Importantly, across independent datasets, the same T cell state is enriched in the circulation of patients with clinically aggressive relapsing disease, in CSF-enriched expanded clonotypes, including a subset with proven Epstein-Barr virus-specificity, and in chronically inflamed lesion rims in end-stage MS. Finally, we demonstrate that blockade of lymphocyte trafficking across the blood-brain barrier with the anti-integrin α4 therapy natalizumab leads to enrichment of this cell state in the circulation, demonstrating a mechanistically concordant effect across distinct high-efficacy MS therapies. Together, our findings support a B cell-dependent, pathogenic T cell state which links peripheral immune activation to CNS-compartmentalised smouldering neuroinflammation across disease stages. Direct therapeutic manipulation of this T cell state may represent a key opportunity for targeting chronic neuroinflammation and facilitate a strategic shift away from broad immune cell ablation.
A bstract Human disease risk emerges from the shared influences of genetics, environment, lifestyle, and concurrent diseases over time, resulting in recurring patterns of susceptibility across conditions. However, most risk prediction models treat diseases as independent outcomes or rely on limited input variables, restricting their ability to capture these shared patterns. Here we present RisQ, a framework that learns a unified representation of human health across diseases, modalities, and time. This representation is queried with natural language to estimate disease risk for arbitrary diseases and prediction horizons. Generalization to unseen disease groups and prediction horizons indicates that information is shared across diseases and time, revealing a common structure of disease risk that is learnable. Trained and validated in 488,170 participants from the UK Biobank and evaluated without retraining in 257,538 participants from the independent All of Us cohort, RisQ leverages this shared structure to outperform disease-specific models, multi-disease frameworks, and tabular foundation models in risk prediction. We show that jointly modeling increasing numbers of diseases, input modalities, and prediction horizons improves performance, indicating that scaling these axes increases information transfer and enriches the learned structure. We then show this structure is multi-scale: it captures demographic determinants of disease susceptibility, while also organizing individuals into reproducible cross-disease risk clusters within demographically restricted subgroups. Genetic analyses further support the biological grounding of the structure by linking gene-level loss of function to cross-disease risk profiles. This surfaces known relationships of HBB , SLC22A12 , CASR , and LDLR , while also highlighting less characterized associations. Together, these results indicate that human disease risk exhibits a shared structure that can be learned from multimodal data to improve risk prediction, stratify individuals by cross-disease susceptibility, and support the discovery of relationships across diseases.
Cell fate transitions are driven by regulatory circuitry, yet RNA velocity models cellular dynamics without explicitly accounting for gene regulatory interactions, limiting mechanistic insight. Conversely, gene regulatory network (GRN) inference methods largely neglect the dynamic nature of biological systems. To overcome this conceptual disconnect, we present RegVelo, a bottom-up, actionable, and interpretable deep learning framework that jointly models splicing kinetics and gene regulatory interactions. Across diverse biological systems, RegVelo provides reliable predictive power for terminal states, gene interactions, and perturbation simulations. By applying RegVelo to zebrafish neural crest development using full-length Smart-seq3 and shared gene expression and chromatin accessibility measurements, we delineate regulatory programs underlying fate specification. Guided by in silico perturbations and validated by CRISPR-Cas9 knockout and single-cell Perturb-seq, we establish tfec as an early driver and elf1 as a regulator of pigment cell fate. RegVelo establishes a quantitative framework for bridging gene regulation and cell fate decisions.
Abstract Genetic prediction of complex phenotypes typically relies on additive linear models, which scale well but cannot capture non-additive effects or deeply integrate molecular and clinical data. Domain-specific neural networks have driven advances in images, text, and other modalities, but genome-scale neural networks remain challenging because genotypes are sparse and high-dimensional, effective sample sizes are limited, and generic architectures lack interpretability. Here, we introduce the omnigenic neural network, a biologically structured architecture inspired by the omnigenic model of complex traits. The model learns hierarchical representations of biological processes, accommodates multimodal inputs, supports transfer learning, and enables multitask prediction. Models trained in the UK Biobank and evaluated in the All of Us cohort for ischemic heart disease, type 2 diabetes, and schizophrenia outperformed published PGS Catalog and PRS-CSx scores. A multitask model trained across 36 cardiovascular endpoints further outperformed corresponding single-phenotype models and baselines. The architecture provides systems-level interpretability by quantifying the contributions of biological processes, which were consistent with established disease mechanisms. It also captures non-linear interactions between variants. Analysis of these interactions using Integrated Hessians revealed patterns concordant with previously reported epistatic associations. Together, these findings establish the omnigenic neural network as a flexible framework for interpretable, multimodal, and multitask genomic prediction.
A central challenge in single-cell biology is distinguishing disease-associated remodeling from normal cellular heterogeneity. Addressing this challenge requires healthy reference frameworks that capture cellular diversity across individuals, technologies, and biological contexts. Here we present the Human Pancreas Cell Atlas (HPCA), a reference atlas of the healthy human pancreas integrating 815,126 single-cell and single-nucleus transcriptomes from 109 donors across 12 studies, diverse technologies, and demographics. Using benchmarked integration and community-driven annotations, HPCA defines 94 cell types and transcriptional states spanning endocrine, exocrine, immune, and stromal compartments. The atlas identifies rare endocrine populations, including a putative, spatially supported polyhormonal alpha-beta-delta state, and provides a unified framework for interpreting pancreatic cellular variation across diverse biological and demographic covariates. Projection of disease and model-system datasets onto HPCA contextualized endocrine and epithelial remodeling relative to healthy pancreatic states. Diabetes-associated endocrine cells remained embedded within the healthy endocrine state space while exhibiting disease-specific changes, as supported by spatial and eQTL concordance analyses. Integration with a pancreatic ductal adenocarcinoma atlas resolved injury-associated and malignant epithelial ecosystem regions across donors. Finally, the HPCA enables quantitative benchmarking of murine diabetes models and stem-cell-derived islets against human pancreatic reference states. Together, the HPCA establishes a healthy transcriptional coordinate system for interpreting disease-associated pathophysiology , experimental perturbation, and regenerative fidelity, illustrating how reference atlases can function as analytical frameworks rather than static cell catalogs.
Cell fate transitions are driven by regulatory circuitry, yet RNA velocity models cellular dynamics without explicitly accounting for gene regulatory interactions, limiting mechanistic insight. Conversely, gene regulatory network (GRN) inference methods largely neglect the dynamic nature of biological systems. To overcome this conceptual disconnect, we present RegVelo, a bottom-up, actionable, and interpretable deep learning framework that jointly models splicing kinetics and gene regulatory interactions. Across diverse biological systems, RegVelo provides reliable predictive power for terminal states, gene interactions, and perturbation simulations. By applying RegVelo to zebrafish neural crest development using full-length Smart-seq3 and shared gene expression and chromatin accessibility measurements, we delineate regulatory programs underlying fate specification. Guided by in silico perturbations and validated by CRISPR-Cas9 knockout and single-cell Perturb-seq, we establish tfec as an early driver and elf1 as a regulator of pigment cell fate. RegVelo establishes a quantitative framework for bridging gene regulation and cell fate decisions.
Myeloid cells, including microglia and perivascular macrophages, are central to Alzheimer's disease (AD) neurobiology, yet their role remains incompletely understood. We profiled 832,505 human myeloid cells from the prefrontal cortex of 1,607 donors spanning the lifespan and showing varying degrees of AD neuropathology. We delineated six subclasses comprising 13 transcriptionally distinct subtypes and identified adaptive changes associated with aging and AD progression. Here we show that a disease-associated microglial subtype, characterized by elevated GPNMB expression and enriched for polygenic AD risk, expands with AD pathology and shows increased phagocytic activity. We identify MITF as an upstream regulator required to maintain this microglial state. Cell-cell interaction analyses prioritize APOE-SORL1 and APOE-TREM2 signaling pairs associated with disease progression. Using human and mouse models, we demonstrate that the neuroprotective effects of this microglial subtype depend on TREM2. These findings provide mechanistic insights into myeloid cell function in aging and AD, aiding therapeutic discovery.
Enhancer-promoter interactions (EPIs) play a central role in gene regulation, but experimental techniques such as Hi-C for mapping these interactions remain costly and labor-intensive. Computational methods have been developed to predict EPIs in silico from DNA sequence and chromatin information; however, there are major challenges with the generalizability and accuracy of predictions by existing methods across cell types and conditions unseen during model training. We developed and validated UniversalEPI, an attention-based deep ensemble model that predicts EPIs up to 2 Mb apart using only DNA sequence and chromatin accessibility (ATAC-seq) data. Unlike models that reconstruct full Hi-C contact maps, UniversalEPI focuses on biologically relevant, sparse chromatin interactions between accessible regulatory elements. It generalizes across both bulk and single-cell ATAC-seq-derived pseudo-bulk datasets, delivering state-of-the-art performance while using fewer input modalities than existing approaches. By modeling predictive uncertainty, UniversalEPI enables statistically robust differential analysis of chromatin interactions across conditions. We demonstrate its utility by tracking dynamic EPIs during human macrophage activation and identifying regulatory differences between cancer cell states in esophageal adenocarcinoma. By providing precalculated Hi-C predictions for 157 ENCODE datasets, UniversalEPI expands the scope and applicability of in silico 3D genome modeling for studying gene regulation in development and disease.
Direct reprogramming of immune cells holds promise for immunotherapy but is constrained by limited knowledge of transcription factor (TF) networks. Here, we developed REPROcode, a combinatorial single-cell screening platform to identify TF combinations for immune cell reprogramming. We first validated REPROcode by inducing type-1 conventional dendritic cells (cDC1s) with multiplexed sets of 9, 22, and 42 factors. With cDC1-enriched TFs, REPROcode enabled identification of optimal TF stoichiometry, fidelity enhancers, and regulators of cDC1 states. We then constructed an arrayed lentiviral library of 408 barcoded immune TFs to explore broader reprogramming capacity. Screening 48 TFs enriched in dendritic cell subsets yielded myeloid and lymphoid phenotypes and enabled the construction of a TF hierarchy map to guide immune reprogramming. Finally, we validated REPROcode’s discovery power by inducing natural killer (NK)-like cells. This study deepens our understanding of immune transcriptional control and provides a versatile toolbox for engineering immune cells to advance immunotherapy.
Abstract Single-cell RNA sequencing (scRNA-seq) profiles transcriptomes at high resolution but discards the spatial context of cells within a tissue—information that is essential for studying intercellular mechanisms and tissue architecture. Spatial transcriptomics (ST) retains coordinates but, depending on the assay, trades this off against gene-panel breadth, spatial resolution, or cost. We present G2T (Gene-to-Tissue), a generative deep learning model that reassembles a tissue from gene expression — its only observed input — by predicting the matrix of pairwise distances between cells in a learned embedding space. G2T uses an attention-based Transformer with an Euclidean-Distance-Matrix (EDM) output head and is trained with conditional flow matching: the network learns to denoise corrupted cell positions, conditioned on the slice’s gene expression, by predicting per-cell embeddings whose pairwise squared distances match the ground-truth distance matrix. At inference, a fast locally-optimal-block (LOBPCG) multidimensional scaling step turns the predicted distance matrix into 2-D coordinates. On a published MERFISH mouse primary motor cortex benchmark, G2T improves over the previous state-of-the-art method, LUNA, across all three standard metrics— Spearman correlation of pairwise-distance ranks, Contact F1, and per-cell-class Sum RSSD — and even larger relative gains on the mouse central-nervous-system scRNA-seq atlas, evaluated against an imputed spatial reference (STARmap PLUS-integrated locations, not measured coordinates). By predicting this geometry in a higher-dimensional embedding space rather than regressing 2-D coordinates, G2T relaxes the 2-D output parameterisation of prior diffusion-based methods and yields a compact, scalable building block for reconstructing tissue from dissociated cells, enabling downstream spatial niche and cell–cell communication analysis.