CRISPR base editing screens are powerful tools for studying disease-associated variants at scale. However, the efficiency and precision of base editing perturbations vary, confounding the assessment of variant-induced phenotypic effects. Here, we provide an integrated pipeline that improves the estimation of variant impact in base editing screens. We perform high-throughput ABE8e-SpRY base editing screens with an integrated reporter construct to measure the editing efficiency and outcomes of each gRNA alongside their phenotypic consequences. We introduce BEAN, a Bayesian network that accounts for per-guide editing outcomes and target site chromatin accessibility to estimate variant impacts. We show this pipeline attains superior performance compared to existing tools in variant classification and effect size quantification. We use BEAN to pinpoint common variants that alter LDL uptake, implicating novel genes. Additionally, through saturation base editing of LDLR, we enable accurate quantitative prediction of the effects of missense variants on LDL-C levels, which aligns with measurements in UK Biobank individuals, and identify structural mechanisms underlying variant pathogenicity. This work provides a widely applicable approach to improve the power of base editor screens for disease-associated variant characterization.
Genome-wide association studies (GWAS) have identified numerous single nucleotide polymorphisms (SNPs) associated with various traits and diseases, yet understanding the functional consequences of these variants remains challenging. We have chosen a set of 18 loci associated with cholesterol traits (LDL-C and HDL-C) in a recent trans-ancestry GWAS (Graham et al 2021, Nature, GLGC consortium). Genes within these loci have coding burden for these same traits and/or are known monogenic disease genes, and importantly, targeting these genes gives robust phenotypes in CRISPR screens using cholesterol-related phenotypic assays. We have used human genetic evidence to select ~2,500 variants within these 18 loci to evaluate, including variants with strong GWAS evidence and variants with strong evidence as liver eQTLs through fine-mapping and/or linkage to sentinel variants. This protocol describes pooled LDL uptake screen using CRISPR base editing screening to install all variants in a human hepatocyte cell line, with the goal of gaining insight into causal variants and genes at the selected 18 loci.
Most current single-cell analysis pipelines are limited to cell embeddings and rely heavily on clustering, while lacking the ability to explicitly model interactions between different feature types. Furthermore, these methods are tailored to specific tasks, as distinct single-cell problems are formulated differently. To address these shortcomings, here we present SIMBA, a graph embedding method that jointly embeds single cells and their defining features, such as genes, chromatin-accessible regions and DNA sequences, into a common latent space. By leveraging the co-embedding of cells and features, SIMBA allows for the study of cellular heterogeneity, clustering-free marker discovery, gene regulation inference, batch effect removal and omics data integration. We show that SIMBA provides a single framework that allows diverse single-cell problems to be formulated in a unified way and thus simplifies the development of new analyses and extension to new single-cell modalities. SIMBA is implemented as a comprehensive Python library ( https://simba-bio.readthedocs.io ).
Recent advances in single-cell omics technologies enable the individual and joint profiling of cellular measurements. Currently, most single-cell analysis pipelines are cluster-centric and cannot explicitly model the interactions between different feature types. In addition, single-cell methods are generally designed for a particular task as distinct single-cell problems are formulated differently. To address these current shortcomings, we present SIMBA , a graph embedding method that jointly embeds single cells and their defining features, such as genes, chromatin accessible regions, and transcription factor binding sequences into a common latent space. By leveraging the co-embedding of cells and features, SIMBA allows for the study of cellular heterogeneity, clustering-free marker discovery, gene regulation inference, batch effect removal, and omics data integration. SIMBA has been extensively applied to scRNA-seq, scATAC-seq, and dual-omics data. We show that SIMBA provides a single framework that allows diverse single-cell analysis problems to be formulated in a unified way and thus simplifies the development of new analyses and integration of other single-cell modalities. SIMBA is implemented as an efficient, comprehensive, and extensible Python library ( https://simba-bio.readthedocs.io ) for the analysis of single-cell omics data using graph embedding.
BACKGROUND:Super-enhancers or stretch enhancers are clusters of active enhancers that often coordinate cell-type specific gene regulation during development and differentiation. In addition, the enrichment of disease-associated single nucleotide polymorphism in super-enhancers indicates their critical function in disease-specific gene regulation. However, little is known about the function of super-enhancers beyond gene regulation.RESULTS:In this study, through a comprehensive analysis of super-enhancers in 30 human cell/tissue types, we identified a new class of super-enhancers which are constitutively active across most cell/tissue types. These 'common' super-enhancers are associated with universally highly expressed genes in contrast to the canonical definition of super-enhancers that assert cell-type specific gene regulation. In addition, the genome sequence of these super-enhancers is highly conserved by evolution and among humans, advocating their universal function in genome regulation. Integrative analysis of 3D chromatin loops demonstrates that, in comparison to the cell-type specific super-enhancers, the cell-type common super-enhancers present a striking association with rapidly recovering loops.CONCLUSIONS:In this study, we propose that a new class of super-enhancers may play an important role in the early establishment of 3D chromatin structure.