Mapping enhancers and their target genes in specific cell types is crucial for understanding gene regulation and human disease genetics. However, accurately predicting enhancer-gene regulatory interactions from single-cell datasets has been challenging. Here we introduce a family of classification models, scE2G, to predict enhancer-gene regulation. These models use features from single-cell assay for transposase-accessible chromatin with sequencing (ATAC-seq) or multiomic RNA and ATAC-seq data, and are trained on a CRISPR perturbation dataset including >10,000 evaluated element-gene pairs. We benchmark scE2G models against CRISPR perturbations, fine-mapped expression quantitative trait loci and genome-wide association study variant-gene associations and demonstrate state-of-the-art performance at prediction tasks across several cell types and categories of perturbations. We apply scE2G to build maps of enhancer-gene regulatory interactions in heterogeneous tissues and interpret noncoding variants associated with complex traits, nominating regulatory interactions linking INPP4B and IL15 to lymphocyte count. The scE2G models will enable accurate mapping of enhancer-gene regulatory interactions across thousands of human cell types.
A major challenge in human genetics is to identify all distal regulatory elements and determine their effects on target gene expression in a given cell type. To this end, large-scale CRISPR screens have been conducted to perturb thousands of candidate enhancers. Using these data, predictive models have been developed that aim to generalize such findings to predict which enhancers regulate which genes across the genome. However, existing CRISPR methods and large-scale datasets have limitations in power, scale, or selection bias, with the potential to skew our understanding of the properties of distal regulatory elements and confound our ability to evaluate predictive models. Here, we develop a new framework for highly powered, unbiased CRISPR screens, including an optimized experimental method (Direct-Capture Targeted Perturb-seq (DC-TAP-seq)), a random design strategy, and a comprehensive analytical pipeline that accounts for statistical power. We applied this framework to survey 1,425 randomly selected candidate regulatory elements across two human cell lines. Our results reveal fundamental properties of distal regulatory elements in the human genome. Most element-gene regulatory interactions are estimated to have small effect sizes (<10%), which previous experiments were not powered to detect. Most cis -regulatory interactions occur over short genomic distances (<100 kb). A large fraction of the discovered regulatory elements bind CTCF but do not show chromatin marks typical of classical enhancers. Housekeeping genes have similar frequencies of distal regulatory elements compared to other genes, but with 2-fold weaker effect sizes. Comparisons to the predictions of the ENCODE-rE2G model suggest that, while performance is similar across two cell types, new models will be needed to detect elements with weaker effect sizes, regulatory effects of CTCF sites, and enhancers for housekeeping genes. Overall, this study describes the first unbiased, perturbation-based survey of thousands of distal regulatory element-gene connections, and provides a framework for expanding such efforts to build more complete maps of distal regulation in the human genome.
Regulatory DNA provides a platform for transcription factor binding to encode cell-type-specific patterns of gene expression. However, the effects and programmability of regulatory DNA sequences remain difficult to map or predict. Here, we develop variant effects from flow-sorting experiments with CRISPR targeting screens (Variant-EFFECTS) to introduce hundreds of designed edits to endogenous regulatory DNA and quantify their effects on gene expression. We systematically dissect and reprogram 3 regulatory elements for 2 genes in 2 cell types. These data reveal endogenous binding sites with effects specific to genomic context, transcription factor motifs with cell-type-specific activities, and limitations of computational models for predicting the effect sizes of variants. We identify small edits that can tune gene expression over a large dynamic range, suggesting new possibilities for prime-editing-based therapeutics targeting regulatory DNA. Variant-EFFECTS provides a generalizable tool to dissect regulatory DNA and to identify genome editing reagents that tune gene expression in an endogenous context.
Linking variants from genome-wide association studies (GWAS) to underlying mechanisms of disease remains a challenge1,4,6. For some diseases, a successful strategy has been to look for cases where multiple GWAS loci contain genes that act in the same biological pathway1–6. However, our knowledge of which genes act in which pathways is incomplete, particularly for cell-type specific pathways or understudied genes. Here we introduce a method to connect GWAS variants to functions, which links variants to genes using epigenomic data, links genes to pathways de novo using Perturb-seq, and integrates these data to identify convergence of GWAS loci onto pathways. We apply this approach to study the role of endothelial cells in genetic risk for coronary artery disease (CAD), and discover that 43 CAD GWAS signals converge on the cerebral cavernous malformations (CCM) signaling pathway. Two regulators of this pathway, CCM2 and TLNRD1,are each linked to a CAD risk variant, regulate other CAD risk genes, and affect atheroprotective processes in endothelial cells. These results suggest a model where CAD risk is driven in part by the convergence of causal genes onto a particular transcriptional pathway in endothelial cells, highlight shared genes between common and rare vascular diseases (CAD and CCM), and identify TLNRD1 as a new, previously uncharacterized member of the CCM signaling pathway. This approach will be widely useful for linking variants to functions for other common polygenic diseases. Note: The list of authors for this protocol does not include all authors of the accompanying manuscript, only those who played a role in developing and executing the Perturb-seq method. For a complete list of manuscript authors, see the Manuscript Citation. Notes are provided to indicate which authors are best to contact for questions regarding specific methods.
Enhancers are key drivers of gene regulation thought to act via 3D physical interactions with the promoters of their target genes. However, genome-wide depletions of architectural proteins such as cohesin result in only limited changes in gene expression, despite a loss of contact domains and loops. Consequently, the role of cohesin and 3D contacts in enhancer function remains debated. Here, we developed CRISPRi of regulatory elements upon degron operation (CRUDO), a novel approach to measure how changes in contact frequency impact enhancer effects on target genes by perturbing enhancers with CRISPRi and measuring gene expression in the presence or absence of cohesin. We systematically perturbed all 1,039 candidate enhancers near five cohesin-dependent genes and identified 34 enhancer-gene regulatory interactions. Of 26 regulatory interactions with sufficient statistical power to evaluate cohesin dependence, 18 show cohesin-dependent effects. A decrease in enhancer-promoter contact frequency upon removal of cohesin is frequently accompanied by a decrease in the regulatory effect of the enhancer on gene expression, consistent with a contact-based model for enhancer function. However, changes in contact frequency and regulatory effects on gene expression vary as a function of distance, with distal enhancers (e.g., >50Kb) experiencing much larger changes than proximal ones (e.g., <50Kb). Because most enhancers are located close to their target genes, these observations can explain how only a small subset of genes - those with strong distal enhancers - are sensitive to cohesin. Together, our results illuminate how 3D contacts, influenced by both cohesin and genomic distance, tune enhancer effects on gene expression.
Linking variants from genome-wide association studies (GWAS) to underlying mechanisms of disease remains a challenge1-3. For some diseases, a successful strategy has been to look for cases in which multiple GWAS loci contain genes that act in the same biological pathway1-6. However, our knowledge of which genes act in which pathways is incomplete, particularly for cell-type-specific pathways or understudied genes. Here we introduce a method to connect GWAS variants to functions. This method links variants to genes using epigenomics data, links genes to pathways de novo using Perturb-seq and integrates these data to identify convergence of GWAS loci onto pathways. We apply this approach to study the role of endothelial cells in genetic risk for coronary artery disease (CAD), and discover 43 CAD GWAS signals that converge on the cerebral cavernous malformation (CCM) signalling pathway. Two regulators of this pathway, CCM2 and TLNRD1, are each linked to a CAD risk variant, regulate other CAD risk genes and affect atheroprotective processes in endothelial cells. These results suggest a model whereby CAD risk is driven in part by the convergence of causal genes onto a particular transcriptional pathway in endothelial cells. They highlight shared genes between common and rare vascular diseases (CAD and CCM), and identify TLNRD1 as a new, previously uncharacterized member of the CCM signalling pathway. This approach will be widely useful for linking variants to functions for other common polygenic diseases.
Abstract Pancreatic cancer (PDAC) is a lethal disease in part because tumor cells exist in distinct transcriptional states (e.g. basal/mesenchymal v.s. classical/epithelial) with unique phenotypic properties that contribute to tumor growth and treatment resistance. Two major mechanisms have been suggested for treatment evasion: (1) the intrinsic resistance of an existing state to a therapy regimen and (2) plasticity of therapy-sensitive states to adopt more resistant states. The relative contribution of these mechanisms to treatment resistance is still poorly understood. Historically, measurement of plasticity in both human patients and mouse models has involved one of three principles: (1) observing a redistribution of cell states in tissue across timepoints or conditions; (2) identifying cells that have genomic, epigenetic or proteomic features of more than one state (mixed states); and (3) performing single-cell cloning of cells and observing the cell states adopted by clonal progeny. While these approaches are observationally consistent with the notion of plasticity, they either fail to definitively prove the existence of plasticity, are restricted in measurements of plasticity outside of native tissues or are unable to quantify the role of plasticity in treatment resistance. Amongst the most well described forms of plasticity in human development and cancer is epithelial-mesenchymal plasticity (EMP), which includes epithelial to mesenchymal transition (EMT) and mesenchymal to epithelial transition (MET). To better understand and quantify the role of EMP in driving treatment resistance of human PDAC, we have developed single-cell multiomic, functional genomic and computational methods applied to patient-derived models and clinical biopsies. We first profiled twelve patient-derived PDAC cell lines by single-cell RNA-seq (scSeq) and learned convergent epithelial and mesenchymal gene programs that were consistent with programs observed in patient samples. We next performed lineage tracing experiments in three PDAC cell lines using an expressed lentiviral barcoding system (ClonMapper). By performing scSeq on these barcoded lines at weekly timepoints over four weeks, we proved the presence of EMP by showing a single cell can produce progeny in both epithelial and mesenchymal states. We next developed a generative probabilistic model of our lineage tracing data. This demonstrated that clones (cells sharing a barcode) had different transition matrices (different EMT and MET rates), thus suggesting each clone has a distinct level of plasticity. Having established this, we focused on identifying genes that might explain the differing plasticity properties of clones. Using elastic net regression we identified 50 transcription factors (TFs) whose expression significantly explained the propensity for EMP over time across clonal populations. Among these were were several known EMP TFs (Zeb1 and Gata6), understudied TFs (Elf3, Sox2, Sox4, Klf3, Klf5 and Atf4) and novel TFs (Meis2, Meis3, FoxA1 and the interferon regulatory factors Irf6, Irf7 and Irf9). Using single-cell multiomics (paired scSeq and single-cell ATAC-seq) on our barcoded population, we found that 9 of the 50 predicted TFs, including Elf3, had differential accessibility between clones with different plasticity properties, suggesting a role for epigenetic regulation of these TFs in facilitating EMP. Importantly, we leveraged our multiomic data to infer gene regulatory networks influenced by these TFs and found an enrichment of binding motifs for these TFs in enhancer regions of genes in epithelial and mesenchymal programs. To study the role of these predicted TFs in modulating EMP, we developed a CRISPRi system that enabled gene perturbations alongside lineage tracing. We performed a CRISPRi perturb-seq experiment (CRISPR perturbation with scSeq readouts), perturbing the 50 predicted TFs above and 10 control genes, and collected 1.5 million single-cell transcriptomic profiles, in addition to two other CRISPR KO perturb-seq experiments. We performed a negative binomial regression to estimate effect sizes of guide RNAs on all genes. 60% of our guides had significant perturbation effects on their target gene. We subsequently found Klf3 as an important regulator of PDAC proliferation independent of cell state. Importantly, we found that knockdown (KD) of several factors, such has Grhl2 and FoxA1, bias towards mesenchymal cell states, whereas KD of others such as Batf2, Snai1, Rel, Zeb1, Nr2f1 and Sox4 led to a bias towards epithelial cell states. Interestingly, KD of several TFs influenced transition properties of cells by decreasing rates of EMT and MET across barcodes without biasing the overall clonal distribution towards a single cell state. This suggests a role for these TFs in enabling plasticity and facilitating state transitions. To study the effect of plasticity in treatment resistance, we treated four barcoded cell lines with the first-line chemotherapy combination FOLFIRINOX (5-fluorouracil, oxaliplatin and SN-38, the active metabolite of irinotecan) or targeted therapy and performed scSeq yielding over 600,00 single-cell transcriptomic profiles. We found an enrichment after treatment with FOLFIRINOX of barcodes that were biased for cells in mesenchymal states, consistent with selection against epithelial cells, but also of those barcodes with the highest inferred state transition rates. With targeted therapies, we found selective depletion of mesenchymal states. We next treated our CRISPR perturbed cell lines and found overall a significantly more restricted barcode diversity in cells containing guide RNAs targeting plasticity factors compared to non-targeting controls, suggestive of the role of plasticity in facilitation resistance. To validate the role of these proposed plasticity factors in human patients, we collected paired biopsy samples from 23 patients in a phase 2 clinical trial of metastatic PDAC patients being treated with radiation therapy and dual checkpoint blockade (NCT03104439). We performed scSeq on these samples, and used a supervised Bayesian matrix factorization approach (Spectra) to learn epithelial and mesenchymal gene programs within tumor cells. We subsequently classified cells as epithelial, mesenchymal and intermediate cell types using a gaussian mixture model on gene expression features. We found the intermediate states were enriched in expression of our proposed plasticity factors, and importantly high expression of these factors in baseline samples correlated with a redistribution of states in follow-up biopsies. Our efforts define a robust experimental and quantitative framework for studying tumor cell plasticity in patient-derived model systems with validation in human patient samples using single-cell and spatial transcriptomics. Collectively, we nominate several regulators that alter the propensity of EMP in PDAC, thus posing a paradigm whereby perturbations may be used to homogenize tumor populations towards treatment-sensitive phenotypes for combination therapy. Citation Format: Arnav Mehta, Lynn Bi, Deepika Yeramosu, Michael Bogaev, Martin Jankowiak, Abigail Collins, Aziz Al'Khafaji, Milan Parikh, Mehrtash Babadi, Kyle Evans, Alex Bloemendal, Russell Kunnes, Marc Schwartz, Glen Munson, Elisa Donnard, Thouis R. Jones, Ben Z. Stanger, Jay Shendure, Jonathan Weissman, David T. Ting, Andrew Aguirre, Nir Hacohen, Dana Pe'er, Eric S. Lander. Dissecting and quantifying pancreatic cancer plasticity using single-cell multiomics, lineage tracing and functional genomics reveals novel mediators of therapy resistance [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2024; Part 2 (Late-Breaking, Clinical Trial, and Invited Abstracts); 2024 Apr 5-10; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2024;84(7_Suppl):Abstract nr NG08.
Regulatory DNA sequences within enhancers and promoters bind transcription factors to encode cell type-specific patterns of gene expression. However, the regulatory effects and programmability of such DNA sequences remain difficult to map or predict because we have lacked scalable methods to precisely edit regulatory DNA and quantify the effects in an endogenous genomic context. Here we present an approach to measure the quantitative effects of hundreds of designed DNA sequence variants on gene expression, by combining pooled CRISPR prime editing with RNA fluorescence in situ hybridization and cell sorting (Variant-FlowFISH). We apply this method to mutagenize and rewrite regulatory DNA sequences in an enhancer and the promoter of PPIF in two immune cell lines. Of 672 variant-cell type pairs, we identify 497 that affect PPIF expression. These variants appear to act through a variety of mechanisms including disruption or optimization of existing transcription factor binding sites, as well as creation of de novo sites. Disrupting a single endogenous transcription factor binding site often led to large changes in expression (up to -40% in the enhancer, and -50% in the promoter). The same variant often had different effects across cell types and states, demonstrating a highly tunable regulatory landscape. We use these data to benchmark performance of sequence-based predictive models of gene regulation, and find that certain types of variants are not accurately predicted by existing models. Finally, we computationally design 185 small sequence variants (≤10 bp) and optimize them for specific effects on expression in silico. 84% of these rationally designed edits showed the intended direction of effect, and some had dramatic effects on expression (-100% to +202%). Variant-FlowFISH thus provides a powerful tool to map the effects of variants and transcription factor binding sites on gene expression, test and improve computational models of gene regulation, and reprogram regulatory DNA.
Systematic evaluation of the impact of genetic variants is critical for the study and treatment of human phys-iology and disease. While specific mutations can be introduced by genome engineering, we still lack scalable approaches that are applicable to the important setting of primary cells, such as blood and immune cells. Here, we describe the development of massively parallel base-editing screens in human hematopoietic stem and progenitor cells. Such approaches enable functional screens for variant effects across any hemato-poietic differentiation state. Moreover, they allow for rich phenotyping through single-cell RNA sequencing readouts and separately for characterization of editing outcomes through pooled single-cell genotyping. We efficiently design improved leukemia immunotherapy approaches, comprehensively identify non-coding variants modulating fetal hemoglobin expression, define mechanisms regulating hematopoietic differentia-tion, and probe the pathogenicity of uncharacterized disease-associated variants. These strategies will advance effective and high-throughput variant-to-function mapping in human hematopoiesis to identify the causes of diverse diseases.
Pancreatic cancer is a lethal disease in part because tumor cells exist in distinct transcriptional phenotypes (e.g. basal and classical states), each with a selective ability to evade current chemotherapy regimens. Two major mechanisms have been suggested for treatment evasion: 1) intrinsic resistance of certain phenotypes to particular chemotherapy regimens and 2) plasticity of treatment sensitive phenotypes to adopt more resistant phenotypes. However, the relative contribution of these mechanisms to treatment resistance is still poorly understood. Whereas previous work has described the redistribution of tumor cell states under selective treatment pressure, there is no direct evidence that tumor cells exhibit phenotypic plasticity at steady state or with treatment. By leveraging technological advancements in single-cell methods, lineage tracing and functional genomics, we have now shown direct evidence of phenotypic state switching in human pancreatic cancer cell lines. By performing single-cell RNA-seq on 5 barcoded PDAC cell lines over a steady state timecourse and under chemotherapy selective pressure (>600k cells total), we identify unique plasticity phenotypes within these cell lines and infer regulators of these plastic states. We validate the role of several of these regulators using bulk phenotypic CRISPRi screens in these cell lines. We next perform CRISPRi perturbations along with lineage tracing and single-cell multiomics (>300k cells) to dissect the regulatory relationships that underlie these cell states. We identify several novel epithelial and mesenchymal biasing factors, including those with unique roles in the most plastic clones. Collectively, we nominate several regulators that bias PDAC cell states thus posing a paradigm whereby perturbations may be used to homogenize tumor populations towards treatment-sensitive phenotypes. We believe this approach combined with current chemotherapy regimens could benefit pancreatic cancer patients by targeting residual, resistant tumor cells in the localized and metastatic disease settings to improve patient survival. Citation Format: Arnav Mehta, Lynn Bi, Aziz Al'Khafaji, Martin Jankowiak, Milan Parikh, Mehrtash Babadi, Alex Bloemendal, Marc Schwartz, Glen Munson, Joeseph Chan, Cassandra Burdziak, Elisa Donnard, Ryan Park, Chen Lu, Philippe Rigollet, Andrew Aguirre, Vidya Subramanian, Ray Jones, Eric S. Lander, David T. Ting, Dana Pe'er, Nir Hacohen. Quantifying and dissecting pancreatic cancer cell phenotypic plasticity using lineage tracing, single-cell multiomics and CRISPR perturbations reveals novel regulators of plastic states [abstract]. In: Proceedings of the AACR Special Conference on Pancreatic Cancer; 2022 Sep 13-16; Boston, MA. Philadelphia (PA): AACR; Cancer Res 2022;82(22 Suppl):Abstract nr B016.
Genome-wide association studies (GWAS) have discovered thousands of risk loci for common, complex diseases, each of which could point to genes and gene programs that influence disease. For some diseases, it has been observed that GWAS signals converge on a smaller number of biological programs, and that this convergence can help to identify causal genes1–6. However, identifying such convergence remains challenging: each GWAS locus can have many candidate genes, each gene might act in one or more possible programs, and it remains unclear which programs might influence disease risk. Here, we developed a new approach to address this challenge, by creating unbiased maps to link disease variants to genes to programs (V2G2P) in a given cell type. We applied this approach to study the role of endothelial cells in the genetics of coronary artery disease (CAD). To link variants to genes, we constructed enhancer-gene maps using the Activity-by-Contact model7,8. To link genes to programs, we applied CRISPRi-Perturb-seq9–12 to knock down all expressed genes within ±500 Kb of 306 CAD GWAS signals13,14 and identify their effects on gene expression programs using single-cell RNA-sequencing. By combining these variant-to-gene and gene-to-program maps, we find that 43 of 306 CAD GWAS signals converge onto 5 gene programs linked to the cerebral cavernous malformations (CCM) pathway—which is known to coordinate transcriptional responses in endothelial cells15, but has not been previously linked to CAD risk. The strongest regulator of these programs is TLNRD1, which we show is a new CAD gene and novel regulator of the CCM pathway. TLNRD1 loss-of-function alters actin organization and barrier function in endothelial cells in vitro, and heart development in zebrafish in vivo. Together, our study identifies convergence of CAD risk loci into prioritized gene programs in endothelial cells, nominates new genes of potential therapeutic relevance for CAD, and demonstrates a generalizable strategy to connect disease variants to functions.
Introduction: Genome-wide association studies (GWAS) have discovered >300 associations for CAD. Few loci are functionally characterized and may represent new mechanisms of disease. Hypothesis: Can a high-throughput, unbiased transcriptional screen for all candidate GWAS genes identify the causal genes and biological pathways at CAD GWAS loci in endothelial cells (ECs)? Methods: We applied CRISPRi-Perturb-seq to knock down the expression of all genes within 500 Kb of coronary artery disease GWAS loci (2,300 genes in total) and measure their effects on the transcriptome using single-cell RNA-seq. Results: We identified 60 programs of co-expressed genes, which represent core cellular pathways (i.e. ribosome biogenesis) and EC-specific pathways such as flow response and angiogenesis. The EC-specific programs show the greatest contribution to CAD heritability. 6 EC-specific programs have the greatest number of CAD GWAS candidate genes. One of the EC-specific programs had genes regulated by KLF -transcription factors and was enriched for shear-stress response genes. The novel CAD candidate genes in this program included multiple known mediators of the cerebral cavernous malformation (CCM)-signaling complex. 9 known members of the CCM-signaling cascade and several potentially novel mediators of CAD clustered together in the KLF Perturb-seq topic. These included genetic variants in the CCM2 , KLF4 , RAC1 , and HEG1 loci as well as genes will a similar transcriptional profile but no prior connection to the CCM-signaling complex. Conclusions: High-throughput functional analysis of 2,300 genes proximal to CAD GWAS loci prioritized pathways—such as angiogenesis and EC migration—that are regulated by multiple risk SNPs. Our study identifies new genes that likely influence risk for CAD, identifies convergence of CAD genes into certain pathways in endothelial cells, and demonstrates a generalizable strategy to connect disease variants to functions.
Genome-wide association studies (GWAS) have identified thousands of noncoding loci that are associated with human diseases and complextraits, each of which could reveal insights into the mechanisms of disease(1). Many ofthe underlying causal variants may affect enhancers(2,3), but we lack accurate maps of enhancers and their target genes to interpret such variants. We recently developed the activity-by-contact (ABC) model to predict which enhancers regulate which genes and validated the model using CRISPR perturbations in several cell types(4). Here we apply this ABC model to create enhancer-gene maps in 131 human cell types and tissues, and use these maps to interpret the functions of GWAS variants. Across 72 diseases and complex traits, ABC links 5,036 GWAS signals to 2,249 unique genes, including a class of 577genesthat appear to influence multiple phenotypes through variants in enhancers that act in different cell types. In inflammatory bowel disease (IBD), causal variants are enriched in predicted enhancers by more than 20-fold in particular cell types such as dendritic cells, and ABC achieves higher precision than other regulatory methods at connecting noncoding variants to target genes. These variant-to-function maps reveal an enhancer that contains an IBD risk variant and that regulates the expression of PPIF to alter the membrane potential of mitochondria in macrophages. Our study reveals principles of genome regulation, identifies genes that affect IBD and provides a resource and generalizable strategy to connect risk variants of common diseases to their molecular and cellular functions.
Genome-wide association studies have now identified tens of thousands of noncoding loci associated with human diseases and complex traits, each of which could reveal insights into biological mechanisms of disease. Many of the underlying causal variants are thought to affect enhancers, but we have lacked genome-wide maps of enhancer-gene regulation to interpret such variants. We previously developed the Activity-by-Contact (ABC) Model to predict enhancer-gene connections and demonstrated that it can accurately predict the results of CRISPR perturbations across several cell types. Here, we apply this ABC Model to create enhancer-gene maps in 131 cell types and tissues, and use these maps to interpret the functions of fine-mapped GWAS variants. For inflammatory bowel disease (IBD), causal variants are >20-fold enriched in enhancers in particular cell types, and ABC outperforms other regulatory methods at connecting noncoding variants to target genes. Across 72 diseases and complex traits, ABC links 5,036 GWAS signals to 2,249 unique genes, including a class of 577 genes that appear to influence multiple phenotypes via variants in enhancers that act in different cell types. Guided by these variant-to-function maps, we show that an enhancer containing an IBD risk variant regulates the expression of PPIF to tune mitochondrial membrane potential. Together, our study reveals insights into principles of genome regulation, illuminates mechanisms that influence IBD, and demonstrates a generalizable strategy to connect common disease risk variants to their molecular and cellular functions.
Mammalian genomes harbor millions of noncoding elements called enhancers that quantitatively regulate gene expression, but it remains unclear which enhancers regulate which genes. Here we describe an experimental approach, based on CRISPR interference, RNA FISH, and flow cytometry (CRISPRi-FlowFISH), to perturb enhancers in the genome, and apply it to test >3,000 potential regulatory enhancer-gene connections across multiple genomic loci. A simple equation based on a mechanistic model for enhancer function performed remarkably well at predicting the complex patterns of regulatory connections we observe in our CRISPR dataset. This Activity-by-Contact (ABC) model involves multiplying measures of enhancer activity and enhancer-promoter 3D contacts, and can predict enhancer-gene connections in a given cell type based on chromatin state maps. Together, CRISPRi-FlowFISH and the ABC model provide a systematic approach to map and predict which enhancers regulate which genes, and will help to interpret the functions of the thousands of disease risk variants in the noncoding genome.
Enhancer elements in the human genome control how genes are expressed in specific cell types and harbor thousands of genetic variants that influence risk for common diseases1–4. Yet, we still do not know how enhancers regulate specific genes, and we lack general rules to predict enhancer–gene connections across cell types5,6. We developed an experimental approach, CRISPRi-FlowFISH, to perturb enhancers in the genome, and we applied it to test >3,500 potential enhancer–gene connections for 30 genes. We found that a simple activity-by-contact model substantially outperformed previous methods at predicting the complex connections in our CRISPR dataset. This activity-by-contact model allows us to construct genome-wide maps of enhancer–gene connections in a given cell type, on the basis of chromatin state measurements. Together, CRISPRi-FlowFISH and the activity-by-contact model provide a systematic approach to map and predict which enhancers regulate which genes, and will help to interpret the functions of the thousands of disease risk variants in the noncoding genome. Combining CRISPRi-FlowFISH to perturb enhancers with an activity-by-contact model to predict complex connections allows systematic mapping of enhancer–gene connections in a given cell type, on the basis of chromatin-state measurements.
Various cis-regulatory functions of genomic loci that produce long non-coding RNAs are revealed, including instances where their promoters have enhancer-like activity and the lncRNA transcripts themselves are not required for activity. Since the discovery of pervasive transcription of long non-coding RNAs (lncRNAs) in mammalian genomes, there has been pressure to determine their functions. Here, Eric Lander and colleagues use a CRISPR/Cas9 deletion approach to uncover various cis-regulatory functions of lncRNAs, including instances in which their promoters have enhancer-like activity and the lncRNA transcripts themselves are often not required for activity. Such effects on neighbouring genes are also seen for protein-coding loci. Mammalian genomes are pervasively transcribed1,2 to produce thousands of long non-coding RNAs (lncRNAs)3,4. A few of these lncRNAs have been shown to recruit regulatory complexes through RNA–protein interactions to influence the expression of nearby genes5,6,7, and it has been suggested that many other lncRNAs can also act as local regulators8,9. Such local functions could explain the observation that lncRNA expression is often correlated with the expression of nearby genes2,10,11. However, these correlations have been challenging to dissect12 and could alternatively result from processes that are not mediated by the lncRNA transcripts themselves. For example, some gene promoters have been proposed to have dual functions as enhancers13,14,15,16, and the process of transcription itself may contribute to gene regulation by recruiting activating factors or remodelling nucleosomes10,17,18. Here we use genetic manipulation in mouse cell lines to dissect 12 genomic loci that produce lncRNAs and find that 5 of these loci influence the expression of a neighbouring gene in cis. Notably, none of these effects requires the specific lncRNA transcripts themselves and instead involves general processes associated with their production, including enhancer-like activity of gene promoters, the process of transcription, and the splicing of the transcript. Furthermore, such effects are not limited to lncRNA loci: we find that four out of six protein-coding loci also influence the expression of a neighbour. These results demonstrate that cross-talk among neighbouring genes is a prevalent phenomenon that can involve multiple mechanisms and cis-regulatory signals, including a role for RNA splice sites. These mechanisms may explain the function and evolution of some genomic loci that produce lncRNAs and broadly contribute to the regulation of both coding and non-coding genes.
Mammalian genomes are pervasively transcribed to produce thousands of spliced long noncoding RNAs (lncRNAs), whose functions remain poorly understood. Because recent evidence has implicated several specific lncRNA loci in the local regulation of gene expression, we sought to determine whether such local regulation is a property of many lncRNA loci. We used genetic manipulations to dissect 12 genomic loci that produce lncRNAs and found that 5 of these loci influence the expression of a neighboring gene in cis . Surprisingly, however, none of these effects required the specific lncRNA transcripts themselves and instead involved general processes associated with their production, including enhancer-like activity of gene promoters, the process of transcription, and the splicing of the transcript. Interestingly, such effects are not limited to lncRNA loci: we found similar effects on local gene expression at 4 of 6 protein-coding loci. These results demonstrate that ‘crosstalk’ among neighboring genes is a prevalent phenomenon that can involve multiple mechanisms and cis regulatory signals, including a novel role for RNA splicing. These mechanisms may explain the function and evolution of some genomic loci that produce lncRNAs.
Gene expression in mammals is regulated by noncoding elements that can affect physiology and disease, yet the functions and target genes of most noncoding elements remain unknown. We present a high-throughput approach that uses clustered regularly interspaced short palindromic repeats (CRISPR) interference (CRISPRi) to discover regulatory elements and identify their target genes. We assess >1 megabase of sequence in the vicinity of two essential transcription factors, MYC and GATA1, and identify nine distal enhancers that control gene expression and cellular proliferation. Quantitative features of chromatin state and chromosome conformation distinguish the seven enhancers that regulate MYC from other elements that do not, suggesting a strategy for predicting enhancer-promoter connectivity. This CRISPRi-based approach can be applied to dissect transcriptional networks and interpret the contributions of noncoding genetic variation to human disease.
Although thousands of large intergenic non-coding RNAs (lincRNAs) have been identified in mammals, few have been functionally characterized, leading to debate about their biological role. To address this, we performed loss-of-function studies on most lincRNAs expressed in mouse embryonic stem (ES) cells and characterized the effects on gene expression. Here we show that knockdown of lincRNAs has major consequences on gene expression patterns, comparable to knockdown of well-known ES cell regulators. Notably, lincRNAs primarily affect gene expression in trans. Knockdown of dozens of lincRNAs causes either exit from the pluripotent state or upregulation of lineage commitment programs. We integrate lincRNAs into the molecular circuitry of ES cells and show that lincRNA genes are regulated by key transcription factors and that lincRNA transcripts bind to multiple chromatin regulatory proteins to affect shared gene expression programs. Together, the results demonstrate that lincRNAs have key roles in the circuitry controlling ES cell state.