Motivation Spatial sequencing technologies enable the single-cell-level study of molecular organization in tissues. Revealing such spatial patterns relies on accurate cell segmentation. In complex tissues with dense cell packing, segmentation based solely on nuclear staining is insufficient for accurate cell boundary detection. This limitation arises because accurate segmentation necessitates the delineation of cell morphology, which is driven by molecular activities such as cytoskeletal dynamics, cell-cell adhesion, and intercellular signaling. Thus, integrating molecular information, including gene or protein expression, has the potential to improve segmentation, but remains computationally challenging.Results To address this, we developed SegJointGene, a deep learning framework that jointly performs cell segmentation and spatial gene prioritization by integrating nuclei-based images with spatial gene or protein expression data. SegJointGene designs an information-entropy-guided convolutional neural network together with a computational information discarding score to identify genes that are important for cell-type-specific segmentation. The model iteratively refines gene prioritization and cell boundaries, producing convergent segmentation results along with prioritized spatial genes or proteins across cell types. We applied and benchmarked SegJointGene on both simulation and real spatial datasets, including spatial transcriptomics from the mouse hippocampus and distinct regions of the whole mouse brain, as well as spatial proteomics data from human tonsil. Across datasets, SegJointGene outperformed existing methods by 5%-20% in accurately assigning molecular signals to cell boundaries. Robustness analyses further demonstrated stable performance across varying gene numbers and imaging resolutions. In addition, the genes prioritized by SegJointGene were enriched for structural, developmental, and synaptic signaling pathways, supporting their relevance to spatial tissue organization.Availability and implementation The source code and data are available at https://github.com/daifengwanglab/segjointgene.
Down syndrome is a genetic condition that causes intellectual disability and is characterized by early-onset delays in motor, cognitive, and language development. The molecular mechanisms underlying these neurodevelopmental impairments remain poorly understood. We used single-nucleus multiomic sequencing to simultaneously profile gene expression and chromatin accessibility in the Down syndrome prefrontal cortex during early postnatal development, a critical period for synaptogenesis, neural maturation, and developmental neuroimmune interactions. Our findings reveal widespread dysregulation of chromatin accessibility and gene expression, with deficits spanning metabolic and synaptic pathways, oligodendrocyte lineage progression, and a pronounced neuroinflammatory signature. We present a molecular atlas of Down syndrome neuropathology at a critical stage of brain development, highlighting convergent neurodevelopmental and neurodegenerative pathways and informing potential targeted therapies for Down syndrome-associated neuroinflammation.
Spatially resolved CRISPR screening in vivo has been limited to small perturbation panels and subsets of protein-coding RNAs. We present Perturb-DBiT, a method for co-sequencing of spatial total RNA whole transcriptomes and single guide RNAs (sgRNAs) on the same tissue section in situ. In a human cancer metastatic colonization model, we applied large (80,000+) sgRNA panels across tumor colonies in multiple consecutive tissue sections alongside their corresponding total RNA transcriptomes. We linked perturbations affecting long noncoding RNA covariation, microRNA-mRNA interactions and distinct amino acid-specific tRNA alterations to tumor migration and growth. By integrating transcriptional pseudotime trajectories, we further observed the impact of perturbations on clonal dynamics and cooperation. In an immune-competent syngeneic mouse model, investigation of the tumor immune microenvironment indicated distinct, synergistic effects on immune infiltration and suppression. Perturb-DBiT provides a spatially resolved comprehensive view of perturbation responses in complex tissues, including small and large RNA regulation, tumor proliferation, migration, metastasis and immune interactions.
MeCP2 is subject to many post-translational modifications, such as phosphorylation, in response to diverse stimuli. Several MeCP2 phosphorylation sites have been identified and studied individually. However, the combinatorial effect of phosphorylation at multiple sites has not been directly examined. Building on previous studies characterizing single mutations of MeCP2 phosphorylation sites, we created a mouse model carrying two mutations, serine 80 to alanine and serine 421 to glutamic acid (S80A;S421E, or A;E), in which the MeCP2 protein phosphorylation is fixed in a state observed in active neurons. A battery of behavioral tests was conducted to characterize the functional outcomes at the whole animal level, followed by two independent single-cell/nuclei level multiome profiling assays to characterize altered molecular pathways, and regulatory pattern biased towards glutamatergic neurons on both transcription and chromatin level. While integrated analysis of the multi-dimensional datasets identified cell-type specific molecular targets regulated by MeCP2 phosphorylation, the spatially resolved transcription profiling offered an orthogonal platform for target validation and provided additional information from spatial perspective. Together, results from this study revealed that the combinatorial effect of MeCP2 S80 and S421 phosphorylation is not merely synergistic, and the regulatory is biased towards glutamatergic excitatory neurons. More importantly, our study is the first to explore the single-cell multiomics profile changes in a MeCP2 phosphor mutant model and link it to the functional outcomes.
INTRODUCTION:The vascular endothelial growth factor (VEGF) signaling family plays a role in neurodegenerative diseases, including Alzheimer's disease (AD). Previous work has shown widespread effects of the members FLT1, FLT4, and VEGFB on AD outcomes. However, these analyses have focused within the non-Hispanic White (NHW) population. NETHODS:The goal of this study was to analyze the effects of the VEGF family in underrepresented populations, leveraging large and diverse bulk RNA sequencing and tandem mass tag-mass spectrometry (TMT-MS) proteomic data. Outcomes included measures of AD pathology and diagnosis. RESULTS:Within underrepresented populations, we replicated previously reported effects of FLT1 and FLT4, whereby higher protein abundance was observed in the AD brain and was associated with higher neuropathology burden. In stratified analyses, these associations were largely consistent across race and ethnicity. DISCUSSION:This multi-omic study on the role of the VEGF family in AD emphasizes the need for more representative studies focused on therapeutic targets for AD. HIGHLIGHTS:Vascular endothelial growth factor (VEGF) genes and proteins were quantified in four different brain regions. Samples included participants from four different populations. Previously observed effects were replicated in diverse populations. This study is the largest multi-omic study of the vascular endothelial growth factor (VEGF) genes among Alzheimer's disease (AD) participants from diverse populations.
Neuropsychiatric disorders lack effective treatments due to a limited understanding of the underlying cellular and molecular mechanisms. To address this, we integrated population-scale single-cell genomics data and analyzed 23 cell-type-level gene regulatory networks across schizophrenia, bipolar disorder, and autism. Our analysis revealed potential druggable transcription factors co-regulating known risk genes that converge into cell-type-specific co-regulated modules. We applied graph neural networks on those modules to prioritize novel risk genes and leveraged them in a network-based drug repurposing framework to identify 220 drug molecules with the potential for targeting specific cell types. We found evidence for 37 of these drugs in reversing disorder-associated transcriptional phenotypes. Additionally, we discovered 335 drug-cell quantitative trait loci (eQTLs), revealing genetic variation's influence on drug target expression at the cell-type level. Our results provide a single-cell network medicine resource that provides potential mechanistic insights for advancing treatment options for neuropsychiatric disorders.
AbstractINTRODUCTIONBasal forebrain cholinergic neurons (BFCNs) are integral to learning, attention, and memory, and are prone to degeneration in Down syndrome (DS), Alzheimer’s disease, and other neurodegenerative diseases. However, the mechanisms that lead to degeneration of these neurons are not known.METHODSSingle-nuclei gene expression and ATAC sequencing were performed on postmortem human basal forebrain from unaffected control and DS tissue samples at 0-2 years of age (n=4 each).RESULTSSequencing analysis of postmortem human basal forebrain identifies gene expression differences in early postnatal DS early in life. Genes encoding proteins associated with energy metabolism pathways, specifically oxidative phosphorylation and glycolysis, and genes encoding antioxidant enzymes are upregulated in DS BFCNs.DISCUSSIONMultiomic analyses reveal that energy metabolism may be disrupted in DS BFCNs by birth. Increased oxidative phosphorylation and the accumulation of reactive oxygen species byproducts may be early contributors to DS BFCN neurodegeneration.
INTRODUCTION:Basal forebrain cholinergic neurons (BFCNs) are integral to learning, attention, and memory, and are prone to degeneration in Down syndrome (DS), Alzheimer's disease, and other neurodegenerative diseases. However, the mechanisms that lead to the degeneration of these neurons are not known. METHODS:Single-nucleus gene expression and Assay for Transposase-Accessible Chromatin (ATAC) sequencing were performed on postmortem human basal forebrain from unaffected control and DS tissue samples at 0-2 years of age (n = 4 each). RESULTS:Sequencing analysis of postmortem human basal forebrain identifies gene expression differences in DS early in life. Genes encoding proteins associated with energy metabolism pathways, specifically oxidative phosphorylation and glycolysis, and genes encoding antioxidant enzymes are upregulated in DS BFCNs. DISCUSSION:Multiomic analyses reveal that energy metabolism may be disrupted in DS BFCNs by birth. Increased oxidative phosphorylation and the accumulation of reactive oxygen species byproducts may be early contributors to DS BFCN neurodegeneration. HIGHLIGHTS:First multiomic gene expression and ATAC analysis of human basal forebrain. Basal forebrain pathology in DS begins by birth. Cell type proportions are altered in early postnatal DS basal forebrain. Gene expression suggests dysregulated energy metabolism in DS BFCNs. Genes encoding oxidative phosphorylation subunits and glycolysis enzymes are dysregulated in DS BFCNs.
Single cells interact continuously to form a cell environment that drives key biological processes. Cells and cell environments are highly dynamic across time and space, fundamentally governed by molecular mechanisms, such as gene expression. Recent sequencing techniques measure single-cell-level gene expression under specific conditions, either temporally or spatially. Using these datasets, emerging works, such as virtual cells, can learn biologically useful representations of individual cells. However, these representations are typically static and overlook the underlying cell environment and its dynamics. To address this, we developed CellTRIP, a multi-agent reinforcement learning method that infers a virtual cell environment to simulate the cell dynamics and interactions underlying given single-cell data. Specifically, cells are modeled as individual agents with dynamic interactions, which can be learned through self-attention mechanisms via reinforcement learning. CellTRIP also applies novel truncated reward boot-strapping and adaptive input rescaling to stabilize training. We can in-silico manipulate any combination of cells and genes in our learned virtual cell environment, predict spatial and/or temporal cell changes, and prioritize corresponding genes at the single-cell level. We applied and benchmarked CellTRIP on various simulated and real gene expression datasets, including recapitulating cellular dynamic processes simulated by gene regulatory networks and stochastic models, imputing spatial organization of mouse cortical cells, predicting developmental gene expression changes after drug treatment in cancer cells, and spatiotemporal reconstruction of Drosophila embryonic development, demonstrating its outperformance and broad applicability. Interactive manipulation of those virtual cell environments, including in-silico perturbation, can prioritize spatial and developmental genes for single-cell-level changes, enabling the generation of new insights into cell dynamics over time and space. CellTRIP is open source as a general tool and available at [github.com/daifengwanglab/CellTRIP][1]. ### Competing Interest Statement The authors have declared no competing interest. NSF, Career 2144475 NIH, RF1MH128695 [1]: http://github.com/daifengwanglab/CellTRIP
Spatially resolved in vivo CRISPR screening integrates gene editing with spatial transcriptomics to examine how genetic perturbations alter gene expression within native tissue environments. However, current methods are limited to small perturbation panels and the detection of a narrow subset of protein-coding RNAs. We present Perturb-DBiT, a distinct and versatile approach for the simultaneous co-sequencing of spatial total RNA whole-transcriptome and single-guide RNAs (sgRNAs), base-by-base, on the same tissue section. This method enables unbiased discovery of how genetic perturbations influence RNA regulation, cellular dynamics, and tissue architecture in situ. Applying Perturb-DBiT to a human cancer metastatic colonization model, we mapped large panels of sgRNAs across tumor colonies in consecutive tissue sections alongside their corresponding total RNA transcriptomes. This revealed novel insights into how perturbations affect long non-coding RNA (lncRNA) co-variation, microRNA-mRNA interactions, and global and distinct tRNA alterations in amino acid metabolism linked to tumor migration and growth. By integrating transcriptional pseudotime trajectories, we further uncovered the impact of perturbations on clonal dynamics and cooperation. In an immune-competent syngeneic mouse model, Perturb-DBiT enabled investigation of genetic perturbations within the tumor immune microenvironment, revealing distinct and synergistic effects on immune infiltration and suppression. Perturb-DBiT provides a spatially resolved comprehensive view of how genetic knockouts influence diverse molecular and cellular responses including small and large RNA regulation, tumor proliferation, migration, metastasis, and immune interactions, offering a panoramic perspective on perturbation responses in complex tissues.
Single-omics approaches often provide a limited view of complex biological systems, whereas multi-omics integration offers a more comprehensive understanding by combining diverse data views. However, integrating heterogeneous data types and interpreting the intricate relationships between biological features-both within and across different data views-remains a bottleneck. To address these challenges, we introduce COSIME (Cooperative Multi-view Integration and Scalable Interpretable Model Explainer). COSIME uses backpropagation of Learnable Optimal Transport (LOT) to deep neural networks, enabling the learning of latent features from multiple views to predict disease phenotypes. In addition, COSIME incorporates Monte Carlo sampling to efficiently estimate Shapley values and Shapley-Taylor indices, enabling the assessment of both feature importance and their pairwise interactions-synergistically or antagonistically-in predicting disease phenotypes. We applied COSIME to both simulated data and real-world datasets, including single-cell transcriptomics, single-cell spatial transcriptomics, epigenomics, and metabolomics, specifically for Alzheimer's disease-related phenotypes. Our results demonstrate that COSIME significantly improves prediction performance while offering enhanced interpretability of feature relationships. For example, we identified that synergistic interactions between microglia and astrocyte genes associated with AD are more likely to be active at the edges of the middle temporal gyrus as indicated by spatial locations. Finally, COSIME is open-source and available for general use.
Oligodendrocytes are the myelinating cells within the central nervous system, but the mechanisms by which transcription factors (TFs) cooperate for gene regulation in oligodendrocytes remain unclear. We introduce coTF-reg, an analytical framework that integrates scRNA-seq and scATAC-seq data to identify cooperative TFs co-regulating the target gene (TG). First, we identify co-binding TF pairs in the same oligodendrocyte-specific regulatory regions. Next, we train a deep learning model to predict each TG expression using the co-binding TFs’ expressions. Shapley interaction scores reveal high interactions between co-binding TF pairs, such as SOX10-TCF12. Validation using oligodendrocyte eQTLs and their eGenes that are regulated by these cooperative TFs show potential regulatory roles for genetic variants. Experimental validation using ChIP-seq data confirms some cooperative TF pairs, such as SOX10-OLIG2. Prediction performance of our models is evaluated through holdout data and additional datasets, and an ablation study is also conducted. The results demonstrate stable and consistent performance. The authors introduce coTF-reg, an analytical framework that integrates scRNA-seq and scATAC-seq data to identify the co-binding transcription factors that regulate the genes cooperatively in oligodendrocytes.
The complexity of Alzheimer's disease (AD) manifests in diverse clinical phenotypes, including cognitive impairment and neuropsychiatric symptoms (NPSs). Such heterogeneity can be attributed to the variations in cellular and molecular mechanisms across individuals, especially at the single-cell level. However, the etiology of these phenotypes remains elusive. To address this, the PsychAD project generated a population-level single-nucleus RNA-seq dataset comprising over 6 million nuclei from the prefrontal cortex of 1,494 individual brains, covering a variety of AD-related phenotypes that capture cognitive impairment, severity of pathological lesions, and the presence of NPSs. Using emerging deep learning approaches, we analyzed the PsychAD data to score single-cell phenotype associations and identified phenotype associate cells (PACs). We also analyzed personalized single-cell functional genomics involving cell type interactions and gene regulatory networks. In particular, we learned latent representations of individual functional genomics (embeddings) and quantified importance scores of cell types, genes, and their interactions for each individual. Finally, we associated genetic variants with the changes of cell type-gene regulatory networks across individuals, i.e., gene regulatory QTLs (grQTLs), providing novel functional genomic insights compared to existing QTLs. We identified ∼1.5 million PACs and compared them across 27 distinct brain cell subclasses, prioritizing cell subpopulations and their expressed genes across various AD phenotypes. This includes the upregulation of a reactive astrocyte subtype with neuroprotective function in AD resilient donors. Additionally, we identified PACs that link multiple phenotypes, including a subpopulation of protoplasmic astrocytes that alter their gene expression and regulation in AD donors with depression. Moreover, our embeddings improved phenotype classifications and revealed potentially novel subtypes and population trajectories for AD progression, cognitive impairment, and NPSs. Our importance scores prioritized personalized functional genomic information and showed significant differences in cell-type regulatory mechanisms. Such information also allowed us to further identify subpopulation-level biological pathways such as ancestries. Overall, our analyses have advanced in uncovering the cellular and molecular mechanisms underlying diverse AD phenotypes, which has the potential to provide valuable insights for identifying novel diagnostic markers and therapeutic targets. Our results are available through open-source web apps, as well as a personalized single-cell functional genomic atlas for AD.
Motivation:Transcription factor (TF) coordination plays a key role in gene regulation via direct and/or indirect protein-protein interactions (PPIs) and co-binding to regulatory elements on DNA. Single-cell technologies facilitate gene expression measurement for individual cells and cell-type identification, yet the connection between TF-TF coordination and target gene (TG) regulation of various cell types remains unclear. Results:To address this, we introduce our innovative computational approach, Network Regression Embeddings (NetREm), to reveal cell-type TF-TF coordination activities for TG regulation. NetREm leverages network-constrained regularization, using prior knowledge of PPIs among TFs, to analyze single-cell gene expression data, uncovering cell-type coordinating TFs and identifying revolutionary TF-TG candidate regulatory network links. NetREm's performance is validated using simulation studies and benchmarked across several datasets in humans, mice, yeast. Further, we showcase NetREm's ability to prioritize valid novel human TF-TF coordination links in 9 peripheral blood mononuclear and 42 immune cell sub-types. We apply NetREm to examine cell-type networks in central and peripheral nerve systems (e.g. neuronal, glial, Schwann cells) and in Alzheimer's disease versus Controls. Top predictions are validated with experimental data from rat, mouse, and human models. Additional functional genomics data helps link genetic variants to our TF-TG regulatory and TF-TF coordination networks. Availability and implementation:https://github.com/SaniyaKhullar/NetREm.
Spatial sequencing technologies enable the study of molecular activities at single-cell resolution. Unlocking the spatial patterns within this data fundamentally relies on the accurate segmentation of individual cells. However, in complex tissues, nuclei staining is often insufficient for delineating full cell boundaries. To address this, we developed SegJointGene, a novel deep learning framework that performs joint cell segmentation and spatial gene prioritization. The model iteratively refines an initial coarse segmentation by learning from local gene expression patterns, guided by an information entropy-based importance score. We benchmarked SegJointGene on diverse spatial transcriptomics and proteomics datasets. It consistently outperformed existing methods in segmentation accuracy, demonstrated robustness to data variability, and successfully identified biologically relevant genes and proteins essential for defining cell identity.
Single-omics approaches often provide a limited perspective on complex biological systems, whereas multi-omics integration enables a more comprehensive understanding by combining diverse data views. However, integrating heterogeneous data types and interpreting complex relationships between biological features—both within and across views—remains a major challenge. Here, to address these challenges, we introduce COSIME (Cooperative Multi-view Integration with a Scalable and Interpretable Model Explainer). COSIME applies the backpropagation of a learnable optimal transport algorithm to deep neural networks, thus enabling the learning of latent features from several views to predict disease phenotypes. It also incorporates Monte Carlo sampling to enable interpretable assessments of both feature importance and pairwise feature interactions for both within and across views. We applied COSIME to both simulated and real-world datasets—including single-cell transcriptomics, spatial transcriptomics, epigenomics and metabolomics—to predict Alzheimer’s disease-related phenotypes. Benchmarking of existing methods demonstrated that COSIME improves prediction accuracy and provides interpretability. For example, it reveals that synergistic interactions between astrocyte and microglia genes associated with Alzheimer’s disease are more likely to localize at the edges of the middle temporal gyrus. Finally, COSIME is also publicly available as an open-source tool. Choi et al. introduce a machine learning model that integrates diverse multi-view data to predict disease phenotypes. The model includes an interpretable explainer that identifies interacting biological features, such as synergistic genes in astrocytes and microglia associated with Alzheimer’s disease.
Peripheral nerve function depends on the proper formation and maintenance of Schwann cells. During development, large diameter axons are enveloped by single, myelinating Schwann cells whereas clusters of small diameter axons are surrounded by nonmyelinating Schwann cells in Remak bundles. The transcription factor SRY-box 10 (SOX10) is required for the development of Schwann cells, but its functions in different neural crest and glial lineages have shown that SOX10 commonly cooperates with other transcription factors to activate cell-specific gene programs. In a previous study, we compared SOX10-bound regulatory elements in Schwann cells and identified a specific enrichment for nuclear receptor motifs at SOX10 binding sites in Schwann cells. NR2F1 and NR2F2 (Coup-TFI/Coup-TFII) are expressed from neural crest through Schwann cell maturity, and knockdown of nuclear receptors Nr2f1 and Nr2f2 in primary Schwann cells downregulated several myelin genes. In this study, we have identified a NR2F-regulated target gene network in Schwann cells, which revealed enrichment for nonmyelinating Schwann cell genes. Cut&Run assays in S16 Schwann cells revealed novel, genome-wide binding sites of NR2F2 and two downstream transcription factors: 1) Retinoid X Receptor Gamma (RXRG), which is also expressed preferentially in nonmyelinating Schwann cells, and 2) TEA-Domain factor (TEAD1), which is an important HIPPO pathway component required for Remak bundle formation. Our study elucidates the transcriptional cooperation that programs the regulatory network in nonmyelinating Schwann cells.
SUMMARY:Cellular processes like development, differentiation, and disease progression are highly complex and dynamic (e.g. gene expression). These processes often undergo cell population changes driven by cell birth, proliferation, and death. Single-cell sequencing enables gene expression measurement at the cellular resolution, allowing us to decipher cellular and molecular dynamics underlying these processes. However, the high costs and destructive nature of sequencing restrict observations to snapshots of unaligned cells at discrete timepoints, limiting our understanding of these processes and complicating the reconstruction of cellular trajectories. To address this challenge, we propose ARTEMIS, a generative model integrating a variational autoencoder (VAE) with unbalanced Diffusion Schrödinger Bridge to model cellular processes by reconstructing cellular trajectories, reveal gene expression dynamics, and recover cell population changes. The VAE maps input time-series single-cell data to a continuous latent space, where trajectories are reconstructed by solving the Schrödinger bridge problem using forward-backward stochastic differential equations (SDEs). A drift function in the SDEs captures deterministic gene expression trends. An additional neural network estimates time-varying kill rates for single cells along trajectories, enabling recovery of cell population changes. Using three scRNA-seq datasets-pancreatic β-cell differentiation, zebrafish embryogenesis, and epithelial-mesenchymal transition (EMT) in cancer cells-we demonstrate that ARTEMIS: (i) outperforms state-of-art methods to predict held-out timepoints, (ii) recovers relative cell population changes over time, and (iii) identifies "drift" genes driving deterministic expression trends in cell trajectories. Furthermore, in silico perturbations show that these genes influence processes like EMT. AVAILABILITY AND IMPLEMENTATION:The code for ARTEMIS: https://github.com/daifengwanglab/ARTEMIS.
The prefrontal cortex (PFC) is critical for myriad high-cognitive functions and is associated with several neuropsychiatric disorders. Here, using Patch-seq and single-nucleus multiomic analyses, we identified genes and regulatory networks governing the maturation of distinct neuronal populations in the PFC of rhesus macaque. We discovered that specific electrophysiological properties exhibited distinct maturational kinetics and identified key genes underlying these properties. We unveiled that RAPGEF4 is important for the maturation of resting membrane potential and inward sodium current in both macaque and human. We demonstrated that knockdown of CHD8, a high-confidence autism risk gene, in human and macaque organotypic slices led to impaired maturation, via downregulation of key genes, including RAPGEF4. Restoring the expression of RAPGEF4 rescued the proper electrophysiological maturation of CHD8-deficient neurons. Our study revealed regulators of neuronal maturation during a critical period of PFC development in primates and implicated such regulators in molecular processes underlying autism.