Environmental exposures play a critical role in shaping physical and mental health, yet integrating such data into biomedical research remains technically complex and fragmented. The EnvironMENTAL Climate, Urbanicity, Environment and Society (CLUES) framework is an open-source, end-to-end workflow for generating individual-level environmental exposure data. CLUES automates the selection and download of open-access geospatial datasets, standardises spatial and temporal formats, and maps projections, and links resulting environmental variables to individual-level biomedical data, requiring no prior expertise in geospatial data. CLUES covers key environmental domains, including urban and natural space, climate and weather extremes, air pollution, and regional socioeconomic conditions. Designed for extensibility and cross-cohort applicability, it enables multidimensional exposure mapping across global settings and adheres to FAIR (Findability, Accessibility, Interoperability and Reusability) and privacy-compliant data protection principles. In this work, we present the CLUES framework and evaluate its scalability, computational performance, and reproducibility for large-scale biomedical research.
The pancreas plays a central role in major human diseases, yet our understanding of its cellular diversity and plasticity remains incomplete. Here, we present a single-cell multiomics atlas of the human pancreas, profiling over four million cells and nuclei from 57 donors across fetal development, adult homeostasis, and type 2 diabetes (T2D). Integrating single-cell RNA sequencing (scRNA-seq)/single-nucleus RNA sequencing (snRNA-seq), snATAC-seq, VASA-seq, spatial transcriptomics (Xenium), and multiplexed proteomics (CODEX), we resolve gene expression, chromatin accessibility, and spatial organization at high resolution. We identify transcriptionally plastic centroacinar-like cells (pCACs) in adults with fetal-like features, delineate endocrine and exocrine lineage trajectories during development, and define HNF1A-defined beta cell epigenetic states. In T2D, we observe shifts in beta cell subtypes and altered regulatory programs. Glucose perturbation of healthy islets reveals cell-type-specific adaptation and stress responses. This atlas provides a foundational framework to understand pancreas biology and the role of cellular plasticity in regeneration and disease.
Abstract Genetic prediction of complex phenotypes typically relies on additive linear models, which scale well but cannot capture non-additive effects or deeply integrate molecular and clinical data. Domain-specific neural networks have driven advances in images, text, and other modalities, but genome-scale neural networks remain challenging because genotypes are sparse and high-dimensional, effective sample sizes are limited, and generic architectures lack interpretability. Here, we introduce the omnigenic neural network, a biologically structured architecture inspired by the omnigenic model of complex traits. The model learns hierarchical representations of biological processes, accommodates multimodal inputs, supports transfer learning, and enables multitask prediction. Models trained in the UK Biobank and evaluated in the All of Us cohort for ischemic heart disease, type 2 diabetes, and schizophrenia outperformed published PGS Catalog and PRS-CSx scores. A multitask model trained across 36 cardiovascular endpoints further outperformed corresponding single-phenotype models and baselines. The architecture provides systems-level interpretability by quantifying the contributions of biological processes, which were consistent with established disease mechanisms. It also captures non-linear interactions between variants. Analysis of these interactions using Integrated Hessians revealed patterns concordant with previously reported epistatic associations. Together, these findings establish the omnigenic neural network as a flexible framework for interpretable, multimodal, and multitask genomic prediction.
Electronic health records (EHRs) offer considerable potential for clinical prediction, but their complexity and heterogeneity challenge traditional machine learning. Domain-specific electronic health record foundation models trained on unlabeled EHR data have shown improved predictive accuracy and generalization. However, their development is constrained by limited data access and site-specific vocabularies. We convert EHR data into plain text by replacing medical codes with natural-language descriptions, enabling general-purpose large language models (LLMs) to produce high-dimensional embeddings for downstream prediction tasks without access to private medical training data. LLM-based embeddings perform on par with a specialized EHR foundation model, CLMBR-T-Base, across 15 clinical tasks from the EHRSHOT benchmark. In an external validation using the UK Biobank, an LLM-based model shows statistically significant improvements for some tasks, which we attribute to higher vocabulary coverage and slightly better generalization. Overall, we reveal a trade-off between the computational efficiency of specialized EHR models and the portability and data independence of LLM-based embeddings.
Imaging-based spatially resolved transcriptomics can localize transcripts within tissue sections in three dimensions. However, cell segmentation, which assigns transcripts to cells, is usually performed in two dimensions and spatial doublets in the vertical dimension result in segmented cells containing transcripts originating from multiple cell types. Here we present a computational tool called ovrlpy that identifies overlapping cells, tissue folds and inaccurate cell segmentation by analyzing transcript localization in three dimensions.
A central challenge in single-cell biology is distinguishing disease-associated remodeling from normal cellular heterogeneity. Addressing this challenge requires healthy reference frameworks that capture cellular diversity across individuals, technologies, and biological contexts. Here we present the Human Pancreas Cell Atlas (HPCA), a reference atlas of the healthy human pancreas integrating 815,126 single-cell and single-nucleus transcriptomes from 109 donors across 12 studies, diverse technologies, and demographics. Using benchmarked integration and community-driven annotations, HPCA defines 94 cell types and transcriptional states spanning endocrine, exocrine, immune, and stromal compartments. The atlas identifies rare endocrine populations, including a putative, spatially supported polyhormonal alpha-beta-delta state, and provides a unified framework for interpreting pancreatic cellular variation across diverse biological and demographic covariates. Projection of disease and model-system datasets onto HPCA contextualized endocrine and epithelial remodeling relative to healthy pancreatic states. Diabetes-associated endocrine cells remained embedded within the healthy endocrine state space while exhibiting disease-specific changes, as supported by spatial and eQTL concordance analyses. Integration with a pancreatic ductal adenocarcinoma atlas resolved injury-associated and malignant epithelial ecosystem regions across donors. Finally, the HPCA enables quantitative benchmarking of murine diabetes models and stem-cell-derived islets against human pancreatic reference states. Together, the HPCA establishes a healthy transcriptional coordinate system for interpreting disease-associated pathophysiology , experimental perturbation, and regenerative fidelity, illustrating how reference atlases can function as analytical frameworks rather than static cell catalogs.
High-grade pancreatic neuroendocrine carcinoma (panNEC) is a rare, aggressive cancer with limited tissue availability. Single-nucleus transcriptomics of five large-cell panNECs revealed two clinically relevant neuroendocrine cell states: one highly proliferative with an aberrant brain-specific neuronal differentiation program and another stress-responsive state enriched for heat stress, hypoxia, and glycolysis. Our findings suggest potential therapeutic vulnerabilities, highlighting the need to evaluate the efficacy of combination therapies. Single-nucleus transcriptomics of large-cell pancreatic neuroendocrine carcinoma uncovers aberrant brain-type neuronal programs and stress-responsive cell states, revealing therapeutic vulnerabilities for this aggressive cancer
BACKGROUND: Drug resistance and lack of predicting biomarkers are a major challenge for cancer therapy. The combination of a BRAF inhibitor (BRAFi) together with an anti-EGFR inhibitor (EGFRi) represents a standard of care approach in BRAFV600E metastatic colorectal cancer (mCRC) patients. However, predictive biomarkers of sensitivity, that could support patient selection for treatment with this combination, are currently missing. Therefore, our goal is to identify those biomarkers associated with response to the combination of BRAFi and EGFRi. METHODS: Here, we established a living biobank of BRAFV600E colorectal cancer patients derived organoids (PDOs) and categorized them as sensitive or resistant to the combination of BRAFi and EGFRi using short term proliferation assays. To elucidate biomarkers of response, drug testing was integrated with genomic, transcriptomic, proteomic and single-cell transcriptmic profiling of our PDOs. RESULTS: Here we revealed the PTEN/PIK3CA/p-AKT axis as mechanism of primary sensitivity while ROS pathway inhibition as driver for primary resistance. Finally, we newly discovered histology and cellular composition as biomarker of drug response. CONCLUSION: These data align with recently published clinical trial data, thus reinforcing the proof that PDOs can be used for biomarker identification. The use of histology and cellular compositions as biomarkers has to be further validated in clinical setting.
ABSTRACT Single‐cell and spatial omics have revolutionized biomedical research by enabling high‐resolution molecular profiling across cells and tissues, thereby overcoming key limitations of bulk sequencing and revealing unprecedented cellular heterogeneity and spatial organization central to development, homeostasis, and disease. Specifically, advances in high‐throughput, subcellular, and multiomics profiling are promoting the field toward deeper insights. In parallel, computational progress, including generative artificial intelligence (AI) and foundation models, is developing rapidly for manipulating multimodal multiomics data. These advancements have been applied to diverse diseases and biological systems, facilitating innovative biomedical findings. However, a significant gap persists between rapid methodological advances and their systematic application for deciphering human biology and pathology. This review synthesizes recent breakthroughs in single‐cell and spatial technologies and surveys computational methods, including AI‐driven approaches, foundation models, and multi‐omics integration algorithms for both single‐cell and spatial analyses. We then summarize representative applications across major human organ systems in health and disease, highlighting opportunities for biomarker discovery, therapeutic target identification, and precision medicine. Finally, we discuss current challenges and future directions for bridging technological innovation with robust biomedical discovery and translational impact. This review provides a vital guide for researchers in the field, offering critical insights for accelerating the translation of single‐cell and spatial omics.
Spatial omics technologies have revolutionized the study of tissue architecture and cellular heterogeneity by integrating molecular profiles with spatial localization. In spatially resolved transcriptomics, delineating higher-order anatomical structures is critical for understanding how cellular organization affects tissue and organ function. Since 2020, more than 50 spatially aware clustering (SAC) methods have been developed for this purpose. However, the reliability of current benchmarks is undermined by their narrow focus on Visium and brain tissue datasets, as well as incorrect interpretation of manual annotation as ground truth. Here, we present SACCELERATOR, a community-driven, extensible framework that standardizes data formatting, method integration, and metric evaluation, and is designed to rapidly incorporate new methods and datasets. SACCELERATOR currently includes 22 SAC methods applied to 15 datasets spanning 9 technologies and diverse tissue types. Our analysis revealed substantial limitations in the generalizability and reproducibility of SAC methods across tissues and platforms. We also demonstrate that anatomical labels commonly used as ground truths are often biased, potentially error-prone, and, in some cases, unsuitable for benchmarking efforts. Rather than scoring and comparing methods, we propose a consensus-guided workflow that aggregates clustering results to generate consensus representations. Descriptive spatial metrics highlight areas of high entropy where method disagreement is highest, enabling targeted feedback for tissue experts. Applied to brain and cancer datasets, this approach uncovered biologically meaningful patterns overlooked by individual methods and manual annotations. Our results underscore the need for iterative, expert-in-the-loop analysis and reveal that traditional evaluation metrics do not always capture the subjective qualities of results. By improving tissue annotation and addressing key benchmarking limitations, SACCELERATOR provides a robust foundation for advancing spatial omics research.
The pancreas plays a central role in major human diseases, yet our understanding of its cellular diversity and plasticity remains incomplete. Here, we present a single-cell multiomics atlas of the human pancreas, profiling over four million cells and nuclei from 57 donors across fetal development, adult homeostasis, and type 2 diabetes (T2D). Integrating sc/snRNA-seq, snATAC-seq, VASA-seq, spatial transcriptomics (Xenium), and multiplexed proteomics (CODEX), we resolve gene expression, chromatin accessibility, and spatial organization at high resolution. We identify transcriptionally plastic centroacinar-like cells (pCACs) in adults with fetal-like features, delineate endocrine and exocrine lineage trajectories during development, and uncover HNF1A-defined beta cell epigenetic states. In T2D, we observe shifts in beta cell subtypes and altered regulatory programs. Glucose perturbation of healthy islets reveals cell-type-specific adaptation and stress responses. This atlas provides a foundational framework to understand pancreas biology and the role of cellular plasticity in regeneration and disease.
SARS-CoV-2 infection leads to extensive host transcriptomic changes, but the role of alternative splicing in shaping the immune response remains underexplored. Here, we present the first application of long-read single-cell RNA sequencing on nasopharyngeal swabs from COVID-19 patients and healthy controls to resolve transcript-level changes across cell types. Our analysis identified major epithelial cell types and pronounced immune infiltration, with cell-type annotations concordant with those from short-read data. By enabling isoform-level resolution, our nanopore sequencing approach revealed cell-type specific alternative splicing, undetectable with short-read sequencing. For example, although gene-level expression of the key immune and apoptosis regulators, IFNAR2 and FAIM , did not differ between COVID-19 patients and healthy controls we identified marked shifts in isoform usage. Between moderate and critical cases, we observed cell-type specific differential transcript usage in the T cell signaling kinase FYN and the immune-regulatory transcription factor IRF2 . As some of these splicing alterations yield functionally distinct isoforms, we hypothesize that alternative splicing modulates immune signaling and apoptosis, fine-tuning the host response to SARS-CoV-2 infection. Our study demonstrates the unique power of long-read single-cell transcriptomics to uncover isoform-resolved regulatory changes, offering novel insights into the role of alternative splicing in shaping immune responses to viral infections. ### Competing Interest Statement The authors have declared no competing interest. EASI-Genomics, 824110 Alexander von Humboldt Stiftung Alliance4Rare Research Network
Giant cell myocarditis (GCM) represents the most fulminant form of inflammatory cardiomyopathy, yet its spatial pathophysiology remains poorly defined. We applied full-transcriptome spatial capture in-situ RNA sequencing to a human GCM explant, generating a comprehensive high-resolution molecular atlas across twelve cardiac localizations. Spatial clustering and cell-type deconvolution delineated a complex cellular landscape dominated by cardiomyocytes, fibroblasts, and immune cells, with pronounced transcriptional heterogeneity along the epicardial-endocardial axis. Myeloid cells localized to inflammatory hotspots and exhibited SPP1- and IL1B-driven activation, whereas lymphoid cells displayed a continuum from IgM- to IgG4-secreting plasma-cell differentiation. Layer-resolved pathway analysis revealed epicardial enrichment of IL-6/JAK/STAT3 and TNF-α/NF-κB signaling, myocardial upregulation of oxidative phosphorylation and myogenesis, and endocardial activation of stress and apoptosis programs. These data uncover a layered immune-metabolic architecture linking epicardial inflammation to myocardial remodeling and endocardial stress, providing a spatial framework for understanding immune-mediated myocardial injury in GCM. ### Competing Interest Statement IAJ has received honoraria from AstraZeneca GmbH, Abbott GmbH, Abiomed Inc., and Biotest, unrelated to the submitted work. BH is inventor on patents that use RNA for diagnosis of myocarditis. Patent protection is in process for MCG for diagnosis and measurement of therapy response in inflammatory cardiomyopathy. BH, UL: Patent protection is in process for cytokines for targeted therapy in inflammatory cardiomyopathy and heart failure. ### Funding Statement Nicolas Musigk, Phillip Suwalski, and Bettina Heidecker received grant funding from the Deutsche Herzstiftung e.V.. Bettina Heidecker is funded by the German Heart Center Foundation. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: This study was approved by the ethics committee of Charite Universitaetsmedizin Berlin (EA4/163/21), and the patient provided informed written consent on participating in this study. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study are available upon reasonable request to the authors.
Relapsed and refractory large B-cell lymphomas (r/r LBCL) remain a therapeutic challenge, particularly after CAR-T cell therapy or bispecific antibodies, where prognosis is dismal with no standard treatment. VIPOR(P) is a biologically guided, multi-targeted regimen combining venetoclax (V), ibrutinib (i), prednisone (P), obinutuzumab (O), and lenalidomide (R) ± polatuzumab vedotin (P). The pivotal phase 1b/2 trial achieved an overall response rate (ORR) of 54% and a complete response (CR) rate of 38%, with preferential activity in ABC-DLBCL and high-grade B-cell lymphomas (HGBCL) harboring MYC and BCL2 rearrangements (Melani et al NEJM 2024). The use of multi-targeted agents prompted concerns about the real-world use of VIPOR(P) due to cumulative toxicity, particularly infections in late-line patients. Moreover, the potential of VIPOR(P) as a bridge-to-CAR-T, especially in chemotherapy-refractory patients, remains unexplored.MethodsWe conducted a retrospective, multicenter study of 56 patients with r/r LBCL treated with VIPOR (77%) or VIPOR(P) (23%) across 18 German and Austrian centers between March 2021 and December 2024 (intention to treat [ITT] cohort). Patients were analyzed in the ITT cohort (n=56), and two subcohorts: i) sustained therapy ≥2 cycles (n=24), and ii) bridge-to-CAR-T (n=17). Response assessment by CT/PET-CT, toxicities per CTCAE v5.0. Whole transcriptome and whole exome sequencing were performed to define cell-of-origin (COO) and genetic subtypes (DLBclass and LymphGen).ResultsWithin the ITT cohort, the median age was 60 years (28–76). Patients were heavily pretreated with a median of 4 prior lines (range 1–12) and poor prognostic features: stage III–IV in 91% and elevated LDH in 93%. Overall, 91% had IPI ≥3 before VIPOR(P), 82% were refractory to the last therapy, and 39% had prior CAR-T cell exposure. Histologies included DLBCL, NOS (57%, thereof 69% ABC, 31% GCB), HGBCL with MYC/BCL2 and/or BCL6 rearrangements (30%), HGBCL, NOS (11%), THRLBCL (2%) and transformed indolent lymphomas (27%, also in other histological subgroups as appropriate).Toxicity exceeded pivotal trial rates, reflecting the late-line, frail population, but was manageable with supportive care. Specifically, hematologic toxicity predominated: neutropenia (grade 3/4 56%), thrombocytopenia (grade 3/4 50%), anemia (grade 3/4 40%), and febrile neutropenia (30%, all grades). Grade 3/4 infections occurred in 25% (4 fatal). Two further patients died of a pulmonary embolism and renal failure. Non-hematologic toxicities included transaminitis (61%), hypokalemia (55%), and nausea (27%) and vomiting (28%).VIPOR(P) induced meaningful responses despite high-risk disease characteristics. ORR/CR rates were 45%/13% (ITT), 61%/22% (sustained), and 50%/0% (bridge-to-CAR-T). Molecular stratification underscored the selectivity of VIPOR(P) for specific LBCL subtypes: Responses were enriched in ABC-DLBCL: ORR 71% (ITT), 80% (sustained), and 100% (bridge-to-CAR-T), while no GCB-DLBCL patients responded. Transformed lymphomas demonstrated few delayed but durable responses exceeding 18 months (ORR: 22%). Among 18 WES-profiled patients, 11 pts had high-risk C2 DLBCLs, with biallelic TP53 inactivation and chemo resistance. A trend toward longer PFS was seen in HGBCL-BCL2 and triple-hit cases as well as transformed lymphoma, while HGBCL-NOS responded poorly.Median follow-up was 6 months; median progression-free survival (PFS) and overall survival (OS) in the ITT cohort were 4 and 5 months, respectively. Twelve-month PFS for imaging-assessed patients was 43% (95% CI: 28-67) (ITT), 55% (35-88) (sustained), and 55% (30-100) (bridge-to-CAR-T), and 12-month OS was 25% (95% CI: 13-46), 42% (22-80), and 42% (20-86), respectively. VIPOR(P) successfully bridged 36% of patients to CAR-T and 9% to allo-HSCT, and 22 patients remained alive at data cutoff.ConclusionsVIPOR(P) demonstrated clinically meaningful activity and a manageable safety profile in r/r LBCL refractory to chemo-immunotherapies. Responses dominated in ABC-DLBCL, aligning with the biologically guided design. Despite higher hematological toxicity observed in more heavily pretreated patients than in the pivotal trial, durable responses in sustained responders and successful bridging of chemo-refractory patients to CAR-T cell therapy highlight the relevance of multi-targeted, chemo-free combination therapies for the treatment of r/r LBCL. *contributed equally as first author
Breast cancer (BC) is the most frequently diagnosed cancer among women and a leading cause of cancer-related mortality globally. Accurate and timely diagnosis is essential for improving patient outcomes. However, traditional histopathological assessments are labor-intensive and subjective, leading to inter-observer variability and diagnostic inconsistencies, especially in resource-limited settings. Furthermore, variability in tissue staining, limited availability of standardized annotated datasets, and subtle morphological patterns complicate the consistent characterization of tumors. Deep learning (DL) has recently emerged as a transformative technology in breast cancer pathology, providing automated and objective solutions for cancer detection, classification, and segmentation from histopathological images. This review systematically evaluates advanced deep learning (DL) architectures, including convolutional neural networks (CNNs), generative adversarial networks (GANs), autoencoders, deep belief networks (DBNs), extreme learning machines (ELMs), and transformer-based models such as Vision Transformers (ViTs) as well as transfer learning, attention-based explainable AI techniques, and multimodal integration to address these diagnostic challenges. Analyzing 199 references, including 182 peer-reviewed studies published between 2014 and 2025 and 17 reputable online sources (websites, databases, etc.), we identify key innovations, limitations, and opportunities for future research. Furthermore, we explore the critical roles of synthetic data augmentation, explainable AI (XAI), and multimodal integration to enhance clinical trust, model interpretability, and diagnostic precision, ultimately facilitating personalized and efficient patient care.
Advanced age is the most important risk factor for severe disease or death from COVID-19, but a thorough mechanistic understanding of the molecular and cellular underpinnings is lacking. Multi-omics analysis of 164 samples from SARS-CoV-2-infected persons aged 1 to 84 years reveals a rewiring of type I interferon (IFN) signaling with a gradual shift from signal transducer and activator of transcription 1 (STAT1) to STAT3 activation in monocytes, CD4+ T cells, and B cells with increasing age. Diversion of IFN signaling is associated with increased expression of inflammatory markers, enhanced release of inflammatory cytokines, and delayed contraction of infection-induced CD4+ T cells. A shift from IFN-responsive germinal center B (GCB) cells toward CD69high GCB and atypical B cells during aging correlates with immunoglobulin (Ig)A production in children, whereas complement-fixing IgG predominates in adults. Our data provide a mechanistic basis for inflammation-prone responses to infections and associated pathology during aging.
Molecular changes underlying the persistent health effects after SARS-CoV-2 infection remain poorly understood. To discern the gene regulatory landscape in the upper respiratory tract of COVID-19 patients, we performed enzymatic DNA methylome and single-cell RNA sequencing in nasal cells of COVID-19 patients ( n = 19, scRNA-seq n = 14) and controls ( n = 14, scRNA-seq n = 10). In addition, we resampled a subset of these patients for transcriptome analyses at 3 ( n = 7) and 12 months ( n = 5) post infection and followed the expression of differentially regulated genes over time. Genome-wide DNA methylation analysis revealed 3112 differentially methylated regions between COVID-19 patients and controls. Hypomethylated regions affected immune regulatory genes, while hypermethylated regions were associated with genes governing ciliary function. These genes were not only downregulated in the acute phase of the disease but sustained repressed up to 12 months post infection in ciliated cells. Validation in an independent cohort collected 6 months post infection ( n = 15) indicated symptom-dependent transcriptional repression of ciliary genes. We therefore propose that hypermethylation observed in the acute phase may exert a long-term effect on gene expression, possibly contributing to post-acute COVID-19 sequelae.
DNA methylation-based classification of (brain) tumors has emerged as a powerful and indispensable diagnostic technique. Initial implementations used methylation microarrays for data generation, while most current classifiers rely on a fixed methylation feature space. This makes them incompatible with other platforms, especially different flavors of DNA sequencing. Here, we describe crossNN, a neural network-based machine learning framework that can accurately classify tumors using sparse methylomes obtained on different platforms and with different epigenome coverage and sequencing depth. It outperforms other deep and conventional machine learning models regarding accuracy and computational requirements while still being explainable. We use crossNN to train a pan-cancer classifier that can discriminate more than 170 tumor types across all organ sites. Validation in more than 5,000 tumors profiled on different platforms, including nanopore and targeted bisulfite sequencing, demonstrates its robustness and scalability with 99.1% and 97.8% precision for the brain tumor and pan-cancer models, respectively.