Abstract Determining the activity of cis-regulatory elements (CREs) is essential for modeling gene regulation and interpreting genetic variation. Yet, current methods often lack the specificity to distinguish active regulation from permissive chromatin, the sensitivity to detect unstable enhancer RNAs, or the scalability required to profile limited input material and primary cells. Here, we introduce nucCAGE, a transcription start site (TSS) assay for profiling nuclear, capped RNAs, and PRIME, a computational framework for identifying active CREs from TSS data. Together, these methods increase sensitivity to low-abundance RNAs and enable robust detection of active regulatory elements across diverse contexts. Across multiple orthogonal functional and genetic benchmarks, including fine-mapped eQTLs, ClinVar variants, GWAS loci, and CRISPRi-tested elements, nucCAGE-derived PRIME predictions achieve superior recall compared to state-of-the-art methods while maintaining strong enrichment for phenotype-associated variation. Applying PRIME to the FANTOM5 dataset yields a comprehensive, cell-type-resolved atlas of active CREs that recapitulates known tissue-trait relationships. We demonstrate how this atlas can be used to nominate causal noncoding variants, linking immune-cell enhancer regulation of SMAD3 to asthma and NCOR2 to premature separation of placenta. Together, nucCAGE and PRIME provide a framework for high-sensitivity genome-wide discovery of active CREs and a resource for variant-to-function studies.
Mapping enhancers and their target genes in specific cell types is crucial for understanding gene regulation and human disease genetics. However, accurately predicting enhancer-gene regulatory interactions from single-cell datasets has been challenging. Here we introduce a family of classification models, scE2G, to predict enhancer-gene regulation. These models use features from single-cell assay for transposase-accessible chromatin with sequencing (ATAC-seq) or multiomic RNA and ATAC-seq data, and are trained on a CRISPR perturbation dataset including >10,000 evaluated element-gene pairs. We benchmark scE2G models against CRISPR perturbations, fine-mapped expression quantitative trait loci and genome-wide association study variant-gene associations and demonstrate state-of-the-art performance at prediction tasks across several cell types and categories of perturbations. We apply scE2G to build maps of enhancer-gene regulatory interactions in heterogeneous tissues and interpret noncoding variants associated with complex traits, nominating regulatory interactions linking INPP4B and IL15 to lymphocyte count. The scE2G models will enable accurate mapping of enhancer-gene regulatory interactions across thousands of human cell types.
Abstract Genome-wide association studies (GWAS) have identified thousands of loci associated with cardiometabolic disease, yet translating these associations into regulatory mechanisms, effector genes, and cellular programs remains a major challenge. A key limitation is that genetic effects are often highly context dependent, varying across cell states and environmental conditions that are difficult to model at scale. Here, we leverage CellGenBank , a population-scale biobank of primary human adipose-derived mesenchymal stem cells (AMSCs), to implement a multi-donor cell village in vitro system to map cardiometabolic disease genetic variation across adipocyte differentiation and metabolic stress conditions. We pooled AMSCs from 118 donors into multiplexed villages, differentiated them toward adipocytes, and profiled chromatin accessibility and gene expression using single-nucleus multiome sequencing under four disease-relevant conditions: basal, elevated free fatty acids, low glucose, and hypoxia. By combining with genetic demultiplexing, we quantify how regulatory element activity, gene expression, and higher-order cellular programs are modulated by both genotype and environmental context. Across conditions, we identify widespread context-specific cis-regulatory effects, including expression and chromatin accessibility quantitative trait loci that are masked in baseline states. Genetic effects frequently converge on coordinated transcriptional programs linked to lipid metabolism, insulin responsiveness, and stress adaptation, enabling the identification of cellular program QTLs that bridge variants, genes, and disease-relevant phenotypes. Integration with cardiometabolic GWAS reveals enhanced colocalization in condition- and state-resolved analyses, highlighting the importance of modeling environmental context to resolve disease mechanisms. Together, our study establishes large-scale adipocyte cell villages as a powerful and generalizable framework to map the context-dependent regulatory architecture of cardiometabolic disease and provides a resource linking human genetic variation to adipocyte cellular programs.
Introduction and Objective: Genetic variation and its interaction with environmental cues are crucial in the etiology of metabolic diseases, which are strongly linked to adipocyte dysfunction and are highly cell state- and context-specific. Methods: Here, we leverage a population-scale biobank (CellGenBank) to conduct pooled natural genetic variation screens in primary human adipocyte villages for single-nucleus transcriptomic and chromatin accessibility profiling under various disease-relevant stimuli. Results: Processing 338k nuclei from 118 donors allowed us to identified key cell states characterized by canonical marker genes, including quiescent (PDGFRA) and proliferative (PDGFRA, CDK1, AURKB) adipose tissue-derived mesenchymal stem cells (AMSCs), structural Wnt-regulated adipose tissue-resident (SWAT; DCN, PLAC9, APOD) cells and adipogenic (ADIPOQ, PLIN1) cells, with distinct transcriptional responses to stimuli. By leveraging individual polygenic risk scores, we identified correlations between disease risk and shifts in cell state proportions. These states have been also mapped to metabolic disease relevant traits using single-cell heritability analysis, with body mass index highly enriched in AMSCs, whereas traits informative of metabolic health show strong enrichment in adipogenic cells and a subpopulation of SWAT cells. Furthermore, genome-wide eQTL mapping revealed hundreds of context- and state-specific eQTLs that were enriched in predicted gene-enhancer/promoter regions, linking these variants to their functional transcripts and chromatin accessibility. Conclusion: Collectively, the human adipocyte village approach coupled with single-nucleus functional genomics enabled the discovery of genetic mechanisms underlying metabolic diseases, highlighting its high potential as a tool in the development of effective therapeutic targets. (*Yi Huang and Joaquin Perez-Schindler contribute to this work equally) Y. Huang: None. J. Perez-Schindler: None. S. Datta: None. B. Min: None. N. Nambrath: None. M. Murali: None. H. Dashti: None. B. Sharma: None. S. SinghPoma: None. P. Kubitz: None. W. Qiu: None. R. Andersson: None. T. Jones: None. M. Claussnitzer: None.
Transcription factor (TF) cooperativity plays a critical role in gene regulation. However, the underlying genomic rules remain unclear, calling for scalable methods to characterize the TF binding site (motif) syntax of regulatory elements. Here, we introduce DeepCompARE, a lightweight model paired with an in silico ablation (ISA) framework for genome-wide analysis of regulatory sequences. Our framework enables precise interpretation of the motif syntax governing chromatin accessibility, enhancer activity, and promoter function. We find that most TF motifs are pairwise independent, indicating a default additive behavior of TFs, and define a cooperativity score to quantify deviations from this baseline. This reveals synergy and redundancy as opposite effects along the same cooperative spectrum. TF redundancy is linked to promoter activity and broad expression, whereas TF synergy is associated with enhancer activity, physical interactions, and cell-type specificity. Our framework provides a quantitative model for TF cooperativity, offering new insights into gene regulatory logic.
Although genome wide association studies (GWAS) in large populations have identified hundreds of variants associated with common diseases such as coronary artery disease (CAD), most disease-associated variants lie within non-coding regions of the genome, rendering it difficult to determine the downstream causal gene and cell type. Here, we performed paired single nucleus gene expression and chromatin accessibility profiling from 44 human coronary arteries. To link disease variants to molecular traits, we developed a meta-map of 88 samples and discovered 11,182 single-cell chromatin accessibility quantitative trait loci (caQTLs). Heritability enrichment analysis and disease variant mapping demonstrated that smooth muscle cells (SMCs) harbor the greatest genetic risk for CAD. To capture the continuum of SMC cell states in disease, we used dynamic single cell caQTL modeling for the first time in tissue to uncover QTLs whose effects are modified by cell state and expand our insight into genetic regulation of heterogenous cell populations. Notably, we identified a variant in the COL4A1/COL4A2 CAD GWAS locus which becomes a caQTL as SMCs de-differentiate by changing a transcription factor binding site for EGR1/2. To unbiasedly prioritize functional candidate genes, we built a genome-wide single cell variant to enhancer to gene (scV2E2G) map for human CAD to link disease variants to causal genes in cell types. Using this approach, we found several hundred genes predicted to be linked to disease variants in different cell types. Next, we performed genome-wide Hi-C in 16 human coronary arteries to build tissue specific maps of chromatin conformation and link disease variants to integrated chromatin hubs and distal target genes. Using this approach, we show that rs4887091 within the ADAMTS7 CAD GWAS locus modulates function of a super chromatin interactome through a change in a CTCF binding site. Finally, we used CRISPR interference to validate a distal gene, AMOTL2, liked to a CAD GWAS locus. Collectively we provide a disease-agnostic framework to translate human genetic findings to identify pathologic cell states and genes driving disease, producing a comprehensive scV2E2G map with genetic and tissue level convergence for future mechanistic and therapeutic studies.
The transcription factor MYC is overexpressed in most cancers, where it drives multiple hallmarks of cancer progression. MYC is known to promote oncogenic transcription by binding to active promoters. In addition, MYC has also been shown to invade distal enhancers when expressed at oncogenic levels, but this enhancer binding has been proposed to have low gene-regulatory potential. Here, we demonstrate that MYC directly regulates enhancer activity to promote cancer type-specific gene programs predictive of poor patient prognosis. MYC induces transcription of enhancer RNA through recruitment of RNA polymerase II (RNAPII), rather than regulating RNAPII pause-release, as is the case at promoters. This process is mediated by MYC-induced H3K9 demethylation and acetylation by GCN5, leading to enhancer-specific BRD4 recruitment through its bromodomains, which facilitates RNAPII recruitment. We propose that MYC drives prognostic cancer type-specific gene programs through induction of an enhancer-specific epigenetic switch, which can be targeted by BET and GCN5 inhibitors.
Mapping enhancers and their target genes in specific cell types is crucial for understanding gene regulation and human disease genetics. However, accurately predicting enhancer-gene regulatory interactions from single-cell datasets has been challenging. Here, we introduce a new family of classification models, scE2G, to predict enhancer-gene regulation. These models use features from single-cell ATAC-seq or multiomic RNA and ATAC-seq data and are trained on a CRISPR perturbation dataset including >10,000 evaluated element-gene pairs. We benchmark scE2G models against CRISPR perturbations, fine-mapped eQTLs, and GWAS variant-gene associations and demonstrate state-of-the-art performance at prediction tasks across multiple cell types and categories of perturbations. We apply scE2G to build maps of enhancer-gene regulatory interactions in heterogeneous tissues and interpret noncoding variants associated with complex traits, nominating regulatory interactions linking INPP4B and IL15 to lymphocyte counts. The scE2G models will enable accurate mapping of enhancer-gene regulatory interactions across thousands of diverse human cell types.
Congenital heart defects (CHD) arise in part due to inherited genetic variants that alter genes and noncoding regulatory elements in the human genome. These variants are thought to act during fetal development to influence the formation of different heart structures. However, identifying the genes, pathways, and cell types that mediate these effects has been challenging due to the immense diversity of cell types involved in heart development as well as the superimposed complexities of interpreting noncoding sequences. As such, understanding the molecular functions of both noncoding and coding variants remains paramount to our fundamental understanding of cardiac development and CHD. Here, we created a gene regulation map of the healthy human fetal heart across developmental time, and applied it to interpret the functions of variants associated with CHD and quantitative cardiac traits. We collected single-cell multiomic data from 734,000 single cells sampled from 41 fetal hearts spanning post-conception weeks 6 to 22, enabling the construction of gene regulation maps in 90 cardiac cell types and states, including rare populations of cardiac conduction cells. Through an unbiased analysis of all 90 cell types, we find that both rare coding variants associated with CHD and common noncoding variants associated with valve traits converge to affect valvular interstitial cells (VICs). VICs are enriched for high expression of known CHD genes previously identified through mapping of rare coding variants. Eight CHD genes, as well as other genes in similar molecular pathways, are linked to common noncoding variants associated with other valve diseases or traits via enhancers in VICs. In addition, certain common noncoding variants impact enhancers with activities highly specific to particular subanatomic structures in the heart, illuminating how such variants can impact specific aspects of heart structure and function. Together, these results implicate new enhancers, genes, and cell types in the genetic etiology of CHD, identify molecular convergence of common noncoding and rare coding variants on VICs, and suggest a more expansive view of the cell types instrumental in genetic risk for CHD, beyond the working cardiomyocyte. This regulatory map of the human fetal heart will provide a foundational resource for understanding cardiac development, interpreting genetic variants associated with heart disease, and discovering targets for cell-type specific therapies.
Harsh environments in poorly perfused tumor regions may select for traits driving cancer aggressiveness. Here, we investigated whether tumor acidosis interacts with driver mutations to exacerbate cancer hallmarks. We adapted mouse organoids from normal pancreatic duct (mN10) and early pancreatic cancer (mP4, KRAS-G12D mutation, ± p53 knockout) from extracellular pH 7.4 to 6.7, representing acidic niches. Viability was increased by acid adaptation, a pattern most apparent in wild-type (WT) p53 organoids, and exacerbated upon return to pH 7.4. This led to increased survival of acid-adapted organoids treated with gemcitabine and/or erlotinib, and, in WT p53 organoids, acid-induced attenuation of drug effects. New genetic variants became dominant during adaptation, yet they were unlikely to be its main drivers. Transcriptional changes induced by acid and drug adaptation differed overall, but acid adaptation increased the expression of gemcitabine resistance genes. Thus, adaptation to acidosis increases cancer cell viability after chemotherapy.
ChIA-PET associations between differentially expressed enhancerRNAs (eRNAs) and promoters of annotated genes.
The transcription factor MYC is overexpressed in most cancers, where it drives multiple hallmarks of cancer progression. MYC is known to promote oncogenic transcription by binding to active promoters. In addition, MYC has also been shown to invade distal enhancers when expressed at oncogenic levels, but this enhancer binding has been proposed to have low gene-regulatory potential. Here, we demonstrate that MYC enhancer binding directly promotes cancer type-specific gene programs predictive of poor patient prognosis. MYC induces transcription of enhancer RNA through recruitment of RNAPII, rather than regulating RNAPII pause-release as is the case at promoters. This is mediated by MYC-induced H3K9 demethylation by KDM3A and acetylation by GCN5, leading to enhancer-specific BRD4 recruitment through its bromodomains, which facilitates RNAPII recruitment. Thus, we propose that MYC drives prognostic cancer type-specific gene programs by promoting RNAPII recruitment to enhancers through induction of an epigenetic switch.
Dysfunction of regulatory elements through genetic variants is a central mechanism in the pathogenesis of disease. To better understand disease etiology, there is consequently a need to understand how DNA encodes regulatory activity. Deep learning methods show great promise for modeling of biomolecular data from DNA sequence but are limited to large input data for training. Here, we develop ChromTransfer, a transfer learning method that uses a pre-trained, cell-type agnostic model of open chromatin regions as a basis for fine-tuning on regulatory sequences. We demonstrate superior performances with ChromTransfer for learning cell-type specific chromatin accessibility from sequence compared to models not informed by a pre-trained model. Importantly, ChromTransfer enables fine-tuning on small input data with minimal decrease in accuracy. We show that ChromTransfer uses sequence features matching binding site sequences of key transcription factors for prediction. Together, these results demonstrate ChromTransfer as a promising tool for learning the regulatory code.
Background:The growing prevalence of Alzheimer's disease (AD) is becoming a global health challenge without effective treatments. Defective mitochondrial function and mitophagy have recently been suggested as etiological factors in AD, in association with abnormalities in components of the autophagic machinery like lysosomes and phagosomes. Several large transcriptomic studies have been performed on different brain regions from AD and healthy patients, and their data represent a vast source of important information that can be utilized to understand this condition. However, large integration analyses of these publicly available data, such as AD RNA-Seq data, are still missing. In addition, large-scale focused analysis on mitophagy, which seems to be relevant for the aetiology of the disease, has not yet been performed.Methods:In this study, publicly available raw RNA-Seq data generated from healthy control and sporadic AD post-mortem human samples of the brain frontal lobe were collected and integrated. Sex-specific differential expression analysis was performed on the combined data set after batch effect correction. From the resulting set of differentially expressed genes, candidate mitophagy-related genes were identified based on their known functional roles in mitophagy, the lysosome, or the phagosome, followed by Protein-Protein Interaction (PPI) and microRNA-mRNA network analysis. The expression changes of candidate genes were further validated in human skin fibroblast and induced pluripotent stem cells (iPSCs)-derived cortical neurons from AD patients and matching healthy controls.Results:From a large dataset (AD: 589; control: 246) based on three different datasets (i.e., ROSMAP, MSBB, & GSE110731), we identified 299 candidate mitophagy-related differentially expressed genes (DEG) in sporadic AD patients (male: 195, female: 188). Among these, the AAA ATPase VCP, the GTPase ARF1, the autophagic vesicle forming protein GABARAPL1 and the cytoskeleton protein actin beta ACTB were selected based on network degrees and existing literature. Changes in their expression were further validated in AD-relevant human in vitro models, which confirmed their down-regulation in AD conditions.Conclusion:Through the joint analysis of multiple publicly available data sets, we identify four differentially expressed key mitophagy-related genes potentially relevant for the pathogenesis of sporadic AD. Changes in expression of these four genes were validated using two AD-relevant human in vitro models, primary human fibroblasts and iPSC-derived neurons. Our results provide foundation for further investigation of these genes as potential biomarkers or disease-modifying pharmacological targets.
Genes located within 100kb from the promoters of differentially expressed long non-coding RNAs.