The genomic cis-regulatory code (CRC) underlies spatiotemporal specificity of gene expression. While sequence-to-function (S2F) models can accurately encode the CRC of transcriptional enhancers, decoding these models into human-interpretable rules remains a major challenge. Here we tackle this challenge in human neural development, for which we generate two new single-cell multiome atlases, one from a human embryo and one from neural tube organoids. We use this comparative framework to robustly extract combinations of transcription factor (TF) binding sites that are necessary and sufficient to design enhancers. As such we extract cis-regulatory rules for dorsal-ventral progenitors, neural crest, mesenchyme and neurons. To enable this, we develop a new strategy and computational package, called TF-MINDI, to embed, cluster, and annotate candidate TF binding sites, and to extract combinatorial rules for each cell type. We evaluate rule-based models in conjunction with blackbox S2F models through simulations, evolutionary comparisons with zebrafish, topic modeling, and enhancer reporter-assays. Our findings show robust and interpretable rule extraction and constitute a step forward in deciphering, explaining, and formalizing the CRC. TF-MINDI is available at: https://github.com/aertslab/TF-MINDI. ### Competing Interest Statement The authors have declared no competing interest. ERCAdG, 101054387 KU Leuven, C14/22/125 Chan Zuckerberg Initiative (United States), DI2-0000000068 Interuniversity BOF programme, IBOF/25/010 Stichting tegen Kanker, 2020-1396, 2024-140 Familjen Erling-Perssons Stiftelse, Human Developmental Cell Atlas grant Knut and Alice Wallenberg Foundation, 2015.0041, 2018.0172, 2018.0220 Swedish Foundation for Strategic Research, SB16-0065 EU Horizon2020 BRAINTIME project, 874606 Fonds Wetenschappelijk Onderzoek, 1191323N
Abstract Glioblastoma (GBM) is characterized by diffuse infiltration into the surrounding brain, which precludes complete surgical resection, the strongest determinant of patient survival. The mechanisms that drive this invasive growth remain incompletely understood. Here we identify a lipid-mediated paracrine signaling axis through which microglia, the resident macrophages of the brain, promote glioma invasion. Integrating single-cell transcriptomics, spatial lipidomics, and functional perturbation across mouse models and human GBM samples, we show that invading tumor cells engage and reprogram microglia via CSF1R–PI3K signaling. This induces a metabolic switch in microglia, leading to the secretion of bioactive lipids, including lysophosphatidylcholines (LPCs) and lysophosphatidic acids (LPAs), which act as pro-invasive cues across GBM subtypes through distinct downstream pathways. Disruption of the microglia–GBM axis, either by inhibiting CSF1R signaling or by blocking lipid mobilization, reduces lipid secretion and suppresses tumor invasion. Targeting downstream LPA–LPAR or YAP/TAZ signaling further constrains invasion in a context-dependent manner. Together, these findings define a lipid-driven signaling circuit that links the tumor microenvironment to glioma invasion and identify therapeutic strategies to limit tumor infiltration and improve surgical resectability.
Comparing organ-specific gene regulatory networks (GRNs) across large evolutionary distances remains a major challenge, particularly when the species under study differ in data resolution. This applies to the GRN controlling compound eye development in insects, which is well characterized in Drosophila melanogaster but less well understood in other lineages. Here, we introduce the marmalade fly Episyrphus balteatus (Syrphidae), which diverged from Drosophila approximately 90 million years ago, as a comparative model to study eye development in Diptera. Using RNA-seq and ATAC-seq datasets, we reconstruct the first eye GRN in Episyrphus. Many genes involved in early Drosophila eye specification and differentiation are also active in Episyrphus. Among these, both species share a set of 22 transcription factors (TFs). The GRN built from these TFs and their DNA-binding motifs displays a high degree of internal connectivity. Link conservation analysis, followed by experimental testing in Drosophila, further identifies the AML1/Runx transcription factor lozenge (lz) as a negative regulator of the retinal determination gene dachshund (dac). Within the eye GRN, the degree of regulatory link conservation varies among genes, and both the number and position of regulatory regions often differ across orthologous loci. These results expand the eye GRN of Diptera and suggest extensive network rewiring over the evolutionary span separating Episyrphus and Drosophila.
Fc receptors mediate antibody effector functions. Immunoglobulin G (IgG), the predominant antibody in circulation and in clinical use, engages diverse Fc gamma (Fcγ) receptors differentially expressed across cell types. Here, we provide a comprehensive overview of Fcγ receptor and neonatal Fc receptor (FcRn) expression in humans, macaques, and mice. This analysis revealed substantial differences in Fcγ receptor diversity, cell-specific expression, and regulatory mechanisms that compromise the translation of mouse and macaque models for antibody research. To improve preclinical modeling, we generated a mouse in which humanized Fcγ receptors (FcγRI/CD64, FcγRIIA/CD32A, FcγRIIB/CD32B, FcγRIIIA/CD16A, and FcγRIIIB/CD16B), expressed under control of human promotors, replace their murine counterparts. This model also incorporates human FcRn to improve antibody pharmacokinetics. Humanization resulted in more faithful Fcγ receptor expression. We validated receptor functionality and demonstrated how cytokines modulate their expression. Together, this cross-species Fcγ receptor atlas and humanized mouse model can improve the preclinical evaluation of antibody-based therapeutics.
Sequence-based deep learning models have become the state of the art for analyzing the genomic regulatory code. Particularly for enhancers, these models excel at deciphering sequence grammar that underlies their activity. To enable end-to-end enhancer modeling and design, we developed a software package called CREsted (cis-regulatory element sequence training, explanation and design). It combines preprocessing and analysis of single-cell assay for transposase-accessible chromatin using sequencing data, modeling chromatin accessibility from sequence, sequence design and downstream analysis to decipher enhancer grammar. We demonstrate CREsted's functionality on a mouse cortex and a human peripheral blood mononuclear cell dataset. Additionally, we use CREsted to compare mesenchymal-like cancer cell states between tumor types, and we investigate different fine-tuning strategies of genomic foundation models within CREsted. Finally, we train a model on a zebrafish development atlas and use this to design and in vivo validate cell-type-specific enhancers. For varying datasets, we demonstrate that CREsted facilitates efficient training and analyses, enabling scrutinization of the enhancer logic and design of synthetic enhancers across tissues and species.
Deciphering the cis-regulatory logic underlying cell type identity remains a key challenge in biology. Single-cell chromatin accessibility (scATAC-seq) atlases enable training of sequence-to-function (S2F) deep learning models to decode enhancer logic. Yet, optimal criteria for constructing training datasets, i.e., the number of cells and ATAC fragments, remain unclear. Moreover, the suitability of different scATAC-seq platforms for such models has not been systematically tested. We introduce HyDrop v2, an improved custom droplet scATAC-seq method, and perform the first benchmark of scATAC-seq platforms focusing on its capacity to train S2F models and its capacity to yield TF footprints in different species. We show that lower fragment counts can be compensated for by increased cell numbers. S2F models trained on custom or commercial data perform comparably in enhancer prediction, sequence explainability, and transcription factor footprinting. We demonstrate that integrating data from different scATAC-seq platforms enables large-scale, cost-efficient atlas construction for deep learning-based regulatory modeling. Generating high-quality training data for machine learning is costly. Here, authors include sequence-to-function modeling in benchmarking of custom and commercial droplet-based scATAC platforms, and release a new Drosophila embryo atlas along with a new mouse cortex atlas, assessed for model interpretability.
Genome-wide association studies (GWAS) have linked more than a hundred non-coding genomic loci to Parkinson's disease (PD) risk. Deciphering their functional impact on gene regulation requires cell type-aware modeling approaches to assess the effects of sequence variation on enhancer function and target gene expression. To address this challenge, we generated a comprehensive matched dataset from 190 human donors (115 controls and 75 PD), comprising long-read whole-genome sequencing alongside single nucleus multiome atlases (snATAC-seq and snRNA-seq for 3.1 and 1.1 million nuclei respectively) of the anterior cingulate cortex and substantia nigra. By integrating chromatin accessibility quantitative trait loci (caQTL), DNA methylation QTL (meQTL), and allele-specific chromatin accessibility (ASCA), we identified 53,841 high-confidence cis-acting genetic variants that modulate cell type-specific enhancer accessibility in one or both brain regions. We further demonstrate that sequence-to-function models can accurately predict the impact of these variants directly from the genomic sequence. Novel explainability approaches allowed stratifying these variants according to their regulatory function, with the majority disrupting specific transcription factor binding sites in a cell type specific manner. Integrating these "enhancer variants" (EV) with eQTL mapping and gene locus modeling linked a subset of EVs to their target genes. Finally, we applied these models to prioritize regulatory variants at known PD GWAS loci, bypassing statistical limitations in rare disease-relevant populations like dopaminergic neurons. All together, we establish a unique resource and new sequence modeling strategies to interpret functional non-coding variation in the human brain.
Gene regulatory changes are considered major drivers of evolutionary innovations, including the cerebellum's expansion during human evolution, yet they remain largely unexplored. In this study, we combined single-nucleus measurements of gene expression and chromatin accessibility from six mammals (human, bonobo, macaque, marmoset, mouse, and opossum) to uncover conserved and diverged regulatory networks in cerebellum development. We identified core regulators of cell identity and developed sequence-based models that revealed conserved regulatory codes. By predicting chromatin accessibility across 240 mammalian species, we reconstructed the evolutionary histories of human cis-regulatory elements, identifying sets associated with positive selection and gene expression changes, including the recent gain of THRB expression in cerebellar progenitor cells. Collectively, our work reveals the shared and mammalian lineage-specific regulatory programs governing cerebellum development.
Understanding and modeling how a single human genome concurrently encodes gene regulatory programs for thousands of cell types remains a central challenge in genomics and machine learning. Most human cell types emerge during embryonic, fetal, and pediatric development which are inaccessible to comprehensive molecular profiling. To circumvent this, we hypothesized that the mismatch in evolutionary rates between cis-acting enhancers that modulate gene expression (fast) and the trans-acting regulatory factors that specify cell types (slow) creates an opportunity for 'evolutionary transfer learning'. Specifically, models trained to predict cell type-specific enhancers in one species should generalize to the orthologous cell types and enhancers of related species. To test this, we generated a single-cell atlas of chromatin accessibility spanning mouse embryonic day 10 (E10) to birth (P0). Using combinatorial indexing1, we profiled 3.9 million nuclei from 36 staged embryos, resolving genome-wide accessibility in 36 cell classes and 140 cell types. We then trained a series of multi-output deep learning models (CREsted2), each addressing limitations of the preceding approach, towards the goal of genome-wide prediction of distal enhancers across major developmental lineages. An 'evolution-naive' model achieved strong performance on heldout peaks, but exhibited two failure modes during genome-wide inference: overprediction at tandem repeats and conflation of promoter and distal enhancer grammars. An 'evolution-aware' model resolved these by regrouping accessible regions based on their retention and functional coherence across mammalian evolution, but failed to generalize across species. Finally, an 'evolution-augmented' model, STEAM (Synteny-aware Transfer learning for Enhancer Activity Modeling), incorporated enhancer orthologs from 241 mammalian genomes (Zoonomia3) in a synteny-supervised manner. This increased the effective data scale by as much as 195-fold, markedly improving generalization across mammals despite greater label noise. We applied STEAM to the genome-wide inference of cell class-specific distal developmental enhancers in humans, mice (HumMus) and 239 additional mammals3 (BabaGanoush), i.e. 32 × 241 = 7,712 genome-wide distal enhancer prediction tracks. Together, our results unify advances in single-cell profiling, deep learning, and comparative genomics into a framework for the evolutionary transfer learning of noncoding regulatory grammars. More broadly, our work supports the view that model organisms and evolutionarily diverse genomes are indispensable resources for accelerating and enhancing the AI-enabled exploration of human biology. Note: An interactive version of this preprint, together with count matrices, CREsted models, prediction tracks, code and reproducible figures, is available at this link (ref 4).
Evidence suggests that alternative RNA splicing (AS) plays a critical role in tumor biology and may contribute to the generation of tumor antigens. Here, we develop a method to detect AS in short-read single-cell 5'-RNA-sequencing data, allowing us to uniquely characterize the heterogeneity and dynamic changes in AS in individual cell types within the tumor microenvironment. We identify numerous splicing events specific to either cancer cells or stromal cell types or for triple-negative versus estrogen receptor-positive breast cancers (BCs). By correlating these splice events with expression of splicing regulators in individual cells, we also identify their potential mediators. For instance, we identify and functionally validate the Epithelial Splicing Regulatory Protein-1 (ESRP1) to drive AS in BCs responding to immune checkpoint blockade (ICB). Prioritization of splicing events based on their likelihood to represent tumor antigens reveals that their aggregated load also correlates with high immune activity in multiple cancers, while also predicting expansion of T cells in BCs receiving ICB and prolonging long-term survival of cancer patients treated with ICB. Collectively, our method provides a framework for analyzing AS in single-cell data and defines a key role for AS in the response to ICB.
This protocol details hydrogel bead generation and barcoding for HyDrop v2. The barcodes are build by split-pool and ligation strategy (3x96). Here, we will create an emulsion of acrylamide monomers in a carrier oil containing TEMED. The monomer droplets will polymerise and form hydrogel beads. Ideally, you have a bead stock of around 3 mL of beads before you barcode, but 2 mL can work as well. It is best to produce the beads in one single run, from a single bead mixture to prevent disparities in sizes and/or primer concentration within the bead from occurring.
Developmental processes display temporal differences across species, leading to divergence in organ size and composition. In the cerebral cortex, neurons of diverse identities are generated sequentially through a temporal patterning mechanism conserved throughout mammals. This corticogenesis process is considerably prolonged in the human species, leading to increased brain size and complexity, but the underlying molecular mechanisms remain largely unknown. Here we found that human cortical progenitors displayed lower levels of fatty acid oxidation than their mouse counterparts, in line with their protracted pattern. Treatments that enhance mitochondrial fatty acid oxidation (FAO) accelerated the development of human cortical organoids, including faster progression of neural progenitor cell fate and precocious generation of late-born neurons and glia. FAO accelerated temporal patterning through increased Acetyl-CoA-dependent protein acetylation, including on specific histone transcriptional marks. Thus, species-specific metabolic rates regulate the turnover of post-translation modifications to set the scale of temporal gene regulatory networks of corticogenesis. ### Competing Interest Statement The authors have declared no competing interest.
Astrocytes shape synapses and circuits, yet human basal ganglia astrocyte diversity is incompletely defined. We built a multimodal atlas by integrating single-nucleus RNA-sequencing and chromatin accessibility with DNA methylation, 3D chromatin conformation, and spatial transcriptomics, then mapped basal ganglia programs onto a whole-brain reference. Astrocytes segregated into three anatomical subgroups spanning striatal gray matter, extra-striatal gray matter, and white matter, with subgroup-biased neurotransmitter transporters and synapse-associated programs consistent with differences in dominant afferent input. Within striatum, dorsal and ventral astrocyte populations aligned with distinct microcircuits and were conserved in nonhuman primates. A deep learning sequence model identified subgroup-associated enhancer code and, when benchmarked against published enhancer-AAV datasets, supported the design of candidate viral tools to target basal ganglia astrocyte programs in vivo. Together, these data define major axes of human astrocyte specialization and provide a framework for cell type-specific dissection of basal ganglia function. ### Competing Interest Statement The authors have declared no competing interest. National Institutes of Health, https://ror.org/01cwqze88, UM1MH130981 Research Foundation Flanders (FWO) PhD fellowship, 1SH6J24N & V428025N
Local protein synthesis is vital for neuronal function, but its dysregulation in neurodegenerative diseases remains poorly defined. Here we applied spatial transcriptomics to adult mouse motor nerve axons and cell bodies to enable subcellular mapping. Among transcripts found in mature axons, the most enriched biological process is protein translation, and localization of translation machinery was confirmed using multiplexed single-molecule spatial transcriptomics combined with immunofluorescence. Amyotrophic lateral sclerosis (ALS)-associated mutations in the RNA-binding protein fused in sarcoma (FUS), which suppress local translation, disrupt the compartment-specific RNA signatures, including components of the translation machinery. In particular, eukaryotic initiation factor 5a (Eif5a), a translation factor involved in elongation and termination, is found to be locally impaired in mutant FUS axons with reduced levels of its active hypusinated form. Axon-specific treatment with polyamine spermidine restores Eif5a hypusination and ameliorates mutant FUS-dependent neuronal defects, including suppression of local protein synthesis. Finally, in vivo spermidine treatment reduces ALS-related toxicity in mutant FUS and TDP-43 Drosophila models, which may have implications for therapy development.
This protocol details the preparation of nuclei, bulk ATAC tagmentation, microfluidics setup, HyDrop run, PCR1 cleanup, library purification, and quality control.
Deciphering cis-regulatory logic underlying cell type identity is a fundamental question in biology. Single-cell chromatin accessibility (scATAC-seq) data has enabled training of sequence-to-function deep learning models allowing decoding of enhancer logic and design of synthetic enhancers. Training such models requires large amounts of high-quality training data across species, organs, development, aging, and disease. To facilitate the cost-effective generation of large scATAC-seq atlases for model training, we developed a new version of the open-source microfluidic system HyDrop with increased sensitivity and scale: HyDrop v2. We generated HyDrop v2 atlases for the mouse cortex and Drosophila embryo development and compared them to atlases generated on commercial platforms. HyDrop v2 data integrates seamlessly with commercially available chromatin accessibility methods (10x Genomics). Differentially accessible regions and motif enrichment across cell types are equivalent between HyDrop-v2 and 10x atlases. Sequence-to-function models trained on either atlas are comparable as well in terms of enhancer predictions, sequence explainability, and transcription factor footprinting. By offering accessible data generation, enhancer models trained on HyDrop-v2 and mixed atlases can contribute to unraveling cell-type specific regulatory elements in health and disease. ### Competing Interest Statement The authors have declared no competing interest.
Animal cell types are defined by differential access to genomic information-a process orchestrated by the combinatorial activity of transcription factors that bind to cis-regulatory elements (CREs) to control gene expression. Changes in these gene regulatory networks (GRNs) underlie the origin and diversification of cell types, yet the regulatory logic and specific GRNs that define cell identities remain poorly resolved across the animal tree of life. Cnidarians, as early-branching metazoans, provide a critical window into the early evolution of cell type-specific genome regulation. Here we profiled chromatin accessibility in 60,000 cells from whole adults and gastrula-stage embryos of the sea anemone Nematostella vectensis. We identified 112,728 putative CREs and quantified their activity across cell types, revealing pervasive combinatorial enhancer usage and distinct promoter architectures. To decode the underlying regulatory grammar, we trained sequence-based models predicting CRE accessibility and used these models to infer cell type similarities that reflect known ontogenetic relationships. By integrating sequence motifs, transcription factor expression and CRE accessibility, we reconstructed the GRNs that define cnidarian cell types. Our results show the regulatory complexity underlying cell differentiation in a morphologically simple animal and highlight conserved principles in animal gene regulation. This work provides a foundation for comparative regulatory genomics to understand the evolutionary emergence of animal cell type diversity.
The classical diagnosis of Parkinsonism is based on motor symptoms that are the consequence of nigrostriatal pathway dysfunction and reduced dopaminergic output. However, a decade prior to the emergence of motor issues, patients frequently experience non-motor symptoms, such as a reduced sense of smell (hyposmia). The cellular and molecular bases for these early defects remain enigmatic. To explore this, we developed a new collection of five fruit fly models of familial Parkinsonism and conducted single-cell RNA sequencing on young brains of these models. Interestingly, cholinergic projection neurons are the most vulnerable cells, and genes associated with presynaptic function are the most deregulated. Additional single nucleus sequencing of three specific brain regions of Parkinson’s disease patients confirms these findings. Indeed, the disturbances lead to early synaptic dysfunction, notably affecting cholinergic olfactory projection neurons crucial for olfactory function in flies. Correcting these defects specifically in olfactory cholinergic interneurons in flies or inducing cholinergic signaling in Parkinson mutant human induced dopaminergic neurons in vitro using nicotine, both rescue age-dependent dopaminergic neuron decline. Hence, our research uncovers that one of the earliest indicators of disease in five different models of familial Parkinsonism is synaptic dysfunction in higher-order cholinergic projection neurons and this contributes to the development of hyposmia. Furthermore, the shared pathways of synaptic failure in these cholinergic neurons ultimately contribute to dopaminergic dysfunction later in life.
Connecting neurons into functional circuits requires the formation, maturation, and plasticity of synapses. While advances have been made in identifying individual genes regulating synapse development, the molecular programs orchestrating their action during circuit integration of neurons remain poorly understood. Here, we take a multiomic approach to reconstruct gene regulatory networks (GRNs), comprising transcription factors (TFs), regulatory regions, and predicted target genes, in hippocampal granule cells (GCs). We find a dynamic gene regulatory code, with early and late postnatal GRNs regulating cell morphogenesis and synapse organization and plasticity, respectively. Our results predict sequential regulations, with early-active TFs delaying the activation of later GRNs and their putative synaptic targets. Using a loss-of-function approach, we identify Bcl6 as a regulator of pre- and postsynaptic structural maturation and synaptic transmission and Smad3 as a modulator of inhibitory synaptic transmission in GCs. Together, these findings highlight the networks of key TFs and target genes orchestrating GC synapse development.
Identifying cell-type-specific enhancers is critical for developing genetic tools to study the mammalian brain. We organized the "Brain Initiative Cell Census Network (BICCN) Challenge: Predicting Functional Cell Type-Specific Enhancers from Cross-Species Multi-Omics" to evaluate machine learning and feature-based methods for nominating enhancer sequences targeting mouse cortical cell types. Methods were assessed using in vivo data from hundreds of adeno-associated virus (AAV)-packaged, retro-orbitally delivered enhancers. Open chromatin was the strongest predictor of functional enhancers, while sequence models improved prediction of non-functional enhancers and identified cell-type-specific transcription factor codes to inform in silico enhancer design. This challenge establishes a benchmark for enhancer prioritization and highlights computational and molecular features critical for identifying functional cortical enhancers, advancing efforts to map and manipulate gene regulation in the mammalian cortex.