Cohesin is required for chromatin loop formation. However, its precise role in regulating gene transcription remains largely debated. Here we investigated the relationship between cohesin and RNA polymerase II (RNAPII) using single-molecule mapping and live-cell imaging methods in human cells. Cohesin-mediated transcriptional loops were highly correlated with those of RNA polymerase II and followed the direction of gene transcription. Depleting RAD21, a subunit of cohesin, resulted in the loss of long-range (>100 kb) loops between distal (super-)enhancers and promoters of cell-type-specific downregulated genes. By contrast, short-range (<50 kb) loops were insensitive to RAD21 depletion and connected genes that are mostly constitutively expressed. This result explains why only a small fraction of genes are affected by the loss of long-range chromatin interactions in cohesin-depleted cells. Remarkably, RAD21 depletion appeared to upregulate genes that were involved in initiating DNA replication and disrupted DNA replication timing. Our results elucidate the multifaceted roles of cohesin in establishing transcriptional loops, preserving long-range chromatin interactions for cell-specific genes and maintaining timely DNA replication. Kim, Wang, Clow and colleagues show that long-range chromatin loops bringing distal enhancers or super-enhancers together with promoters are cohesin dependent and cell type specific, whereas most short-range and promoter-centric transcriptional loops are cohesin independent and constitutive.
Dysregulation of the alternative splicing process results in aberrant mRNA transcripts, leading to dysfunctional proteins or nonsense-mediated decay that cause a wide range of mis-splicing diseases. Development of therapeutic strategies to target the alternative splicing process could potentially shift the mRNA splicing from disease isoforms to a normal isoform and restore functional protein. As a proof of concept, we focus on Stargardt disease (STGD1), an autosomal recessive inherited retinal disease caused by biallelic genetic variants in the ABCA4 gene. The splicing variants c.5461-10T>C and c.4773+3A>G in ABCA4 cause the skipping of exon 39-40 and exon 33-34 respectively. In this study, we compared the efficacy of different RNA-targeting systems to modulate these ABCA4 splicing defects, including four CRISPR-Cas13 systems (CASFx-1, CASFx-3, RBFOX1N-dCas13e-C and RBFOX1N-dPspCas13b-C) as well as an engineered U1 system (ExSpeU1). Using a minigene system containing ABCA4 variants in the human retinal pigment epithelium ARPE19, our results show that RBFOX1N-dPspCas13b-C is the best performing CRISPR-Cas system, which enabled up to 80% reduction of the mis-spliced ABCA4 c.5461-10T>C variants and up to 78% reduction of the ABCA4 c.4773+3A>G variants. In comparison, delivery of a single ExSpeU1 was able to effectively reduce the mis-spliced ABCA4 c.4773+3A>G variants by up to 84%. We observed that the effectiveness of CRISPR-based and U1 splicing regulation is strongly dependent on the sgRNA/snRNA targeting sequences, highlighting that optimal sgRNA/snRNA designing is crucial for efficient targeting of mis-spliced transcripts. Overall, our study demonstrated the potential of using RNA-targeting CRISPR-Cas technology and engineered U1 to reduce mis-spliced transcripts for ABCA4 , providing an important step to advance the development of gene therapy to treat STGD1.
Cohesin is required for chromatin loop formation. However, its precise role in regulating gene transcription remains largely unknown. We investigated the relationship between cohesin and RNA Polymerase II (RNAPII) using single-molecule mapping and live-cell imaging methods in human cells. Cohesin-mediated transcriptional loops were highly correlated with those of RNAPII and followed the direction of gene transcription. Depleting RAD21, a subunit of cohesin, resulted in the loss of long-range (>100 kb) loops between distal (super-)enhancers and promoters of cell-type-specific genes. By contrast, the short-range (<50 kb) loops were insensitive to RAD21 depletion and connected genes that are mostly housekeeping. This result explains why only a small fraction of genes are affected by the loss of long-range chromatin interactions due to cohesin depletion. Remarkably, RAD21 depletion appeared to up-regulate genes located in early initiation zones (EIZ) of DNA replication, and the EIZ signals were amplified drastically without RAD21. Our results revealed new mechanistic insights of cohesin's multifaceted roles in establishing transcriptional loops, preserving long-range chromatin interactions for cell-specific genes, and maintaining timely order of DNA replication.
Dictionary learning (DL), implemented via matrix factorization (MF), is commonly used in computational biology to tackle ubiquitous clustering problems. The method is favored due to its conceptual simplicity and relatively low computational complexity. However, DL algorithms produce results that lack interpretability in terms of real biological data. Additionally, they are not optimized for graph-structured data and hence often fail to handle them in a scalable manner.In order to address these limitations, we propose a novel DL algorithm calledonline convex network dictionary learning(online cvxNDL). Unlike classical DL algorithms, online cvxNDL is implemented via MF and designed to handle extremely large datasets by virtue of its online nature. Importantly, it enables the interpretation of dictionary elements, which serve as cluster representatives, through convex combinations of real measurements. Moreover, the algorithm can be applied to data with a network structure by incorporating specialized subnetwork sampling techniques.To demonstrate the utility of our approach, we apply cvxNDL on 3D-genome RNAPII ChIA-Drop data with the goal of identifying important long-range interaction patterns (long-range dictionary elements). ChIA-Drop probes higher-order interactions, and produces data in the form of hypergraphs whose nodes represent genomic fragments. The hyperedges represent observed physical contacts. Our hypergraph model analysis has the objective of creating an interpretable dictionary of long-range interaction patterns that accurately represent global chromatin physical contact maps. Through the use of dictionary information, one can also associate the contact maps with RNA transcripts and infer cellular functions.To accomplish the task at hand, we focus on RNAPII-enriched ChIA-Drop data fromDrosophila MelanogasterS2 cell lines. Our results offer two key insights. First, we demonstrate that online cvxNDL retains the accuracy of classical DL (MF) methods while simultaneously ensuring unique interpretability and scalability. Second, we identify distinct collections of proximal and distal interaction patterns involving chromatin elements shared by related processes across different chromosomes, as well as patterns unique to specific chromosomes. To associate the dictionary elements with biological properties of the corresponding chromatin regions, we employ Gene Ontology (GO) enrichment analysis and perform multiple RNA coexpression studies.
The unique virus-cell interaction in Epstein-Barr virus (EBV)-associated malignancies implies targeting the viral latent-lytic switch is a promising therapeutic strategy. However, the lack of specific and efficient therapeutic agents to induce lytic cycle in these cancers is a major challenge facing clinical implementation. We develop a synthetic transcriptional activator that specifically activates endogenous BZLF1 and efficiently induces lytic reactivation in EBV-positive cancer cells. A lipid nanoparticle encapsulating nucleoside-modified mRNA which encodes a BZLF1-specific transcriptional activator (mTZ3-LNP) is synthesized for EBV-targeted therapy. Compared with conventional chemical inducers, mTZ3-LNP more efficiently activates EBV lytic gene expression in EBV-associated epithelial cancers. Here we show the potency and safety of treatment with mTZ3-LNP to suppress tumor growth in EBV-positive cancer models. The combination of mTZ3-LNP and ganciclovir yields highly selective cytotoxic effects of mRNA-based lytic induction therapy against EBV-positive tumor cells, indicating the potential of mRNA nanomedicine in the treatment of EBV-associated epithelial cancers.
A complete knockout of a single key pluripotency gene may drastically affect embryonic stem cell function and epigenetic reprogramming. In contrast, elimination of only one allele of a single pluripotency gene is mostly considered harmless to the cell. To understand whether complex haploinsufficiency exists in pluripotent cells, we simultaneously eliminated a single allele in different combinations of two pluripotency genes (i.e., Nanog+/-;Sall4+/-, Nanog+/-;Utf1+/-, Nanog+/-;Esrrb+/- and Sox2+/-;Sall4+/-). Although these double heterozygous mutant lines similarly contribute to chimeras, fibroblasts derived from these systems show a significant decrease in their ability to induce pluripotency. Tracing the stochastic expression of Sall4 and Nanog at early phases of reprogramming could not explain the seen delay or blockage. Further exploration identifies abnormal methylation around pluripotent and developmental genes in the double heterozygous mutant fibroblasts, which could be rescued by hypomethylating agent or high OSKM levels. This study emphasizes the importance of maintaining two intact alleles for pluripotency induction.
RNA processing and metabolism are subjected to precise regulation in the cell to ensure integrity and functions of RNA. Though targeted RNA engineering has become feasible with the discovery and engineering of the CRISPR-Cas13 system, simultaneous modulation of different RNA processing steps remains unavailable. In addition, off-target events resulting from effectors fused with dCas13 limit its application. Here we developed a novel platform, Combinatorial RNA Engineering via Scaffold Tagged gRNA (CREST), which can simultaneously execute multiple RNA modulation functions on different RNA targets. In CREST, RNA scaffolds are appended to the 3' end of Cas13 gRNA and their cognate RNA binding proteins are fused with enzymatic domains for manipulation. Taking RNA alternative splicing, A-to-G and C-to-U base editing as examples, we developed bifunctional and tri-functional CREST systems for simultaneously RNA manipulation. Furthermore, by fusing two split fragments of the deaminase domain of ADAR2 to dCas13 and/or PUFc respectively, we reconstituted its enzyme activity at target sites. This split design can reduce nearly 99% of off-target events otherwise induced by a full-length effector. The flexibility of the CREST framework will enrich the transcriptome engineering toolbox for the study of RNA biology.
Three-dimensional (3D) structures of the genome are dynamic, heterogeneous and functionally important. Live cell imaging has become the leading method for chromatin dynamics tracking. However, existing CRISPR- and TALE-based genomic labeling techniques have been hampered by laborious protocols and are ineffective in labeling non-repetitive sequences. Here, we report a versatile CRISPR/Casilio-based imaging method that allows for a nonrepetitive genomic locus to be labeled using one guide RNA. We construct Casilio dual-color probes to visualize the dynamic interactions of DNA elements in single live cells in the presence or absence of the cohesin subunit RAD21. Using a three-color palette, we track the dynamic 3D locations of multiple reference points along a chromatin loop. Casilio imaging reveals intercellular heterogeneity and interallelic asynchrony in chromatin interaction dynamics, underscoring the importance of studying genome structures in 4D.
ABSTRACT Motivation Genomes of multicellular systems are compartmentalized and dynamically folded within the three-dimensional (3D) confines of the nucleus in order to facilitate gene regulation. Among the 3D-genome mapping technologies currently in use, droplet-based, barcode-linked sequencing (ChIA-Drop) has the unique capability to capture complex multi-way chromatin interactions at the single-molecule level. ChIA-Drop data gives rise to higher-order interaction networks in which nodes represent genomic fragments while (hyper)edges capture observed physical contacts. The problem of interest is to use this data to create a “dictionary” of interaction patterns (subnetworks) that accurately describe all global chromatin structures and associate dictionary elements with cellular functions. Results To construct interpretable chromatin dictionaries, we introduce a new algorithm termed online convex network dictionary learning (online cvxNDL). Unlike classical dictionary learning for image or text processing, online cvxNDL uses special subgraph sampling methods and produces interpretable subnetwork representatives corresponding to “convex mixtures” of patterns observed in real data. To demonstrate the utility of the method, we perform an in-depth study of RNAPII-enriched ChIA-Drop data from Drosophila Melanogaster S2 cell lines. Our results are two-fold: First, we show that online cvxNDL allows for accurate reconstruction of the original interaction network data using only a collection of roughly 25 dictionary elements and their “representatives” directly observed in the data. Second, we identify collections of interaction patterns of chromatin elements shared by related processes on different chromosomes and those unique to certain chromosomes. This is accomplished through Gene Ontology (GO) enrichment analysis that allows us to associate dictionary element representatives with functional properties of their corresponding chromatin region and in the process, determine what we call the “span” and “density” of chromatin interaction patterns. Availability and Implementation The code and dataset are available at: https://github.com/jianhao2016/online_cvxNDL/ Contact milenkov@illinois.edu
Zinc finger protein-, transcription activator like effector-, and CRISPR-based methods for genome and epigenome editing and imaging have provided powerful tools to investigate functions of genomes. Targeting sequence design is vital to the success of these experiments. Although existing design software mainly focus on designing target sequence for specific elements, we report here the implementation of Jackie and Albert's Comprehensive K-mer Instances Enumerator (JACKIE), a suite of software for enumerating all single- and multicopy sites in the genome that can be incorporated for genome-scale designs as well as loaded onto genome browsers alongside other tracks for convenient web-based graphic-user-interface-enabled design. We also implement fast algorithms to identify sequence neighborhoods or off-target counts of targeting sequences so that designs with low probability of off-target can be identified among millions of design sequences in reasonable time. We demonstrate the application of JACKIE-designed CRISPR site clusters for genome imaging.
ABSTRACTZFP-, TALE-, and CRISPR-based methods for genome, epigenome editing and imaging have provided powerful tools to interrogate functions of genomes. Targeting sequence design is vital to the success of these experiments. While existing design software mainly focus on designing target sequence for specific elements, we report here the implementation of JACKIE (Jackie and Albert’s Comprehensive K-mer Instances Enumerator), a suite of software for enumerating all single- and multi-copy sites in the genome that can be incorporated for genome-scale designs as well as loaded onto genome browsers alongside other tracks for convenient web-based graphic-user-interface (GUI)-enabled design. We also implement fast algorithms to identify sequence neighborhoods or off-target counts of targeting sequences so that designs with low probability of off-target can be identified among millions of design sequences in reasonable time. We demonstrate the application of JACKIE-designed CRISPR site clusters for genome imaging.
Targeted insertion of exogenous sequences to genomes is useful for therapeutics and biological research. While CRISPR/Cas technologies have been very efficient in gene knockouts by double-strand breaks (DSBs) followed by indel formation through non-homologous end-joining (NHEJ) repair pathway, the precise introduction of new sequences mainly rely on inefficient homology directed repair (HDR) pathways following Cas9-induced DSBs and are restricted to dividing cells. The recent invention of Prime Editing allows short sequences to be precisely inserted at target sites without DSBs. Here, we combine Prime Editing and sequence-specific recombinases and integrases to insert kilobase sequences directionally at target sites. This technique, called PRIMAS for P rime editing, R ecombinase, Integrase- m ediated A ddition of S equence, will expand our genome editing toolbox for targeted insertion of long sequences up to kilobases and beyond.
Abstract Oncogenic extrachromosomal DNA elements (ecDNA) play an important role in tumor evolution, but our understanding of ecDNA biology is limited. We determined the distribution of single-cell ecDNA copy number across patient tissues and cell line models and observed how cell-to-cell ecDNA frequency varies greatly. The exceptional intratumoral heterogeneity of ecDNA suggested ecDNA-specific replication and propagation mechanisms. To evaluate the transfer of ecDNA genetic material from parental to offspring cells during mitosis, we established the CRISPR-based ecTag method. ecTag leverages ecDNA-specific breakpoint sequences to tag ecDNA with fluorescent markers in living cells. Applying ecTag during mitosis revealed disjointed ecDNA inheritance patterns, enabling rapid ecDNA accumulation in individual cells. After mitosis, ecDNAs clustered into ecDNA hubs, and ecDNA hubs colocalized with RNA polymerase II, promoting transcription of cargo oncogenes. Our observations provide direct evidence for uneven segregation of ecDNA and shed new light on mechanisms through which ecDNAs contribute to oncogenesis. Significance: ecDNAs are vehicles for oncogene amplification. The circular nature of ecDNA affords unique properties, such as mobility and ecDNA-specific replication and segregation behavior. We uncovered fundamental ecDNA properties by tracking ecDNAs in live cells, highlighting uneven and random segregation and ecDNA hubs that drive cargo gene transcription. See related commentary by Henssen, p. 293. This article is highlighted in the In This Issue feature, p. 275
CRISPR-Cas technologies enable precise editing of genomic sequences. One major way to introduce precise editing is through homology directed repair (HDR) of DNA double strand breaks (DSB) templated by exogenously supplied single-stranded oligodeoxyribonucleotides (ssODN). Competing pathways determine the outcome of edits. Non-homologous end-joining pathways produce destructive insertions/deletions (indels) at target sites and are dominant over the precise homology directed repair pathways. In this study, we aim to favor HDR and use two strategies to recruit DNA repair proteins (DRPs) to Cas9 cut site, the Casilio-DRP approach that recruits RNA-binding protein-tethered DRPs to target site via aptamers appended to guide RNA; and the 53BP1-DRP approach that recruits DRPs to DSBs via DSB-sensing activity of 53BP1. We conducted two screens using these approaches and identified DRPs such as FANCF and BRCA1 that when recruited to Cas9 cut site, enhance ssODN-templated HDR and increase the proportion of precise edits. This study provides not only new constructs for enhanced CRISPR-ssODN-HDR but also a collection of DRP fusions for studying DNA repair processes. ### Competing Interest Statement The authors have filed a patent application on the invention.
ABSTRACTA complete knockout (KO) of a single key pluripotency gene has been shown to drastically affect embryonic stem cell (ESC) function and epigenetic reprogramming. However, knockin (KI)/KO of a reporter gene only in one of two alleles in a single pluripotency gene is considered harmless and is largely used in the stem cell field. Here, we sought to understand the impact of simultaneous elimination of a single allele in two ESC key genes on pluripotency potential and acquisition. We established multiple pluripotency systems harboring KI/KO in a single allele of two different pluripotency genes (i.e. Nanog+/-; Sall4+/-, Nanog+/-; Utf1+/-, Nanog+/-; Esrrb+/- and Sox2+/-; Sall4+/-). Interestingly, although these double heterozygous mutant lines maintain their stemness and contribute to chimeras equally to their parental control cells, fibroblasts derived from these systems show a significant reduction in their capability to induce pluripotency either by Oct4, Sox2, Klf4 and Myc (OSKM) or by nuclear transfer (NT). Tracing the expression of Sall4 and Nanog, as representative key pluripotency targeted genes, at early phases of reprogramming could not explain the seen delay/blockage. Further exploration identifies abnormal methylation landscape around pluripotent and developmental genes in the double heterozygous mutant fibroblasts. Accordingly, treatment with 5-azacytidine two days prior to transgene induction rescues the reprogramming defects. This study emphasizes the importance of maintaining two intact alleles for pluripotency induction and suggests that insufficient levels of key pluripotency genes leads to DNA methylation abnormalities in the derived-somatic cells later on in development.
Here we describe TALE.Sense, a versatile platform for sensing DNA sequences in live mammalian cells enabling programmable generation of a customable response that discerns cells containing specified sequence targets. The platform is based on the programmable DNA binding of transcription activator-like effector (TALE) coupled to conditional intein-reconstitution producing a trans-spliced ON- switch for a response circuit. TALE.Sense shows higher efficiency and dynamic range when compared to the reported zinc- finger based DNA-sensor in detecting same DNA sequences. Swapping transcriptional activation modules and introducing SunTag-based amplification loops to TALE.Sense circuits augment detection efficiency of the DNA sensor. The TALE.Sense platform shows versatility when applied to a range of target sites, indicating its suitability for applications to identify live cell variants with anticipated DNA sequences. TALE.Sense could be integrated with other cellular or synthetic circuits by using specified DNA sequences as control-switches, thus expanding the scope in connecting inducible modules for synthetic biology.
Abstract Alternative RNA splicing (AS) is a critical step in gene expression regulation. This differential inclusion of sequences in the mRNA is a source of functional transcriptomic diversity that allows multiple RNA isoforms to be expressed from the same gene. Dysregulation of AS is hallmark of cancer and leads to the expression of RNA isoforms that promote tumor initiation and maintenance. Understanding how tumor-associated RNA isoforms impact cancer biology is needed for the development of cancer therapies targeting AS. Here, we describe two approaches to modulate AS in cancer cells. First, we utilize CRISPR-Artificial Splicing Factors (CASFx), splicing factor domains fused to catalytically inactive Cas13 proteins, to direct splicing activity to a specific site using a guide RNA (gRNA). We demonstrate that this approach can be used to promote either exon inclusion or skipping in mRNA from several cancer-associated genes (e.g., HRAS, EP400, TRA2β, SRSF3) depending on the gRNA binding location and the splicing-factor domain fusion. Second, we develop antisense oligonucleotides (ASOs) that target a non-coding poison exon in the oncogenic splicing factor TRA2β. This exon functions as a regulatory “poison exon”, which induces RNA degradation when included, and is differentially spliced in multiple tumor types compared to their corresponding adjacent normal tissue. TRA2β-targeting ASOs promoting poison exon inclusion cause decreased TRA2β protein expression, increased cell death, and decreased proliferation across multiple cancer cell lines. Additionally, compared to cancer cells, normal cell lines were less affected by TRA2β-targeting ASO toxicity, demonstrating a therapeutic window for using this technology to specifically target cancer cells. Citation Format: Nathan K. Leclair, Laura Urbanski, Mattia Brugiolo, Marina Yurieva, Brenton R. Graveley, Albert Cheng, Olga Anczukow. RNA-targeting approaches to modulate alternative RNA splicing in cancer [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2021; 2021 Apr 10-15 and May 17-21. Philadelphia (PA): AACR; Cancer Res 2021;81(13_Suppl):Abstract nr 1147.
The RNA isoform repertoire is regulated by splicing factor (SF) expression, and alterations in SF levels are associated with disease. SFs contain ultraconserved poison exon (PE) sequences that exhibit greater identity across species than nearby coding exons, but their physiological role and molecular regulation is incompletely understood. We show that PEs in serine-arginine-rich (SR) proteins, a family of 14 essential SFs, are differentially spliced during induced pluripotent stem cell (iPSC) differentiation and in tumors versus normal tissues. We uncover an extensive cross-regulatory network of SR proteins controlling their expression via alternative splicing coupled to nonsense-mediated decay. We define sequences that regulate PE inclusion and protein expression of the oncogenic SF TRA2β using an RNA-targeting CRISPR screen. We demonstrate location dependency of RS domain activity on regulation of TRA2β-PE using CRISPR artificial SFs. Finally, we develop splice-switching antisense oligonucleotides to reverse the increased skipping of TRA2β-PE detected in breast tumors, altering breast cancer cell viability, proliferation, and migration.
Chromatin interaction studies can reveal how the genome is organized into spatially confined sub-compartments in the nucleus. However, accurately identifying sub-compartments from chromatin interaction data remains a challenge in computational biology. Here, we present Sub-Compartment Identifier (SCI), an algorithm that uses graph embedding followed by unsupervised learning to predict sub-compartments using Hi-C chromatin interaction data. We find that the network topological centrality and clustering performance of SCI sub-compartment predictions are superior to those of hidden Markov model (HMM) sub-compartment predictions. Moreover, using orthogonal Chromatin Interaction Analysis by in-situ Paired-End Tag Sequencing (ChIA-PET) data, we confirmed that SCI sub-compartment prediction outperforms HMM. We show that SCI-predicted sub-compartments have distinct epigenetic marks, transcriptional activities, and transcription factor enrichment. Moreover, we present a deep neural network to predict sub-compartments using epigenome, replication timing, and sequence data. Our neural network predicts more accurate sub-compartment predictions when SCI-determined sub-compartments are used as labels for training.