Calling Cards is a platform technology to record a cumulative history of transient protein-DNA interactions in the genome of genetically targeted cell types. The record of these interactions is recovered by next generation sequencing. Compared to other genomic assays, whose readout provides a snapshot at the time of harvest, Calling Cards enables correlation of historical molecular states to eventual outcomes or phenotypes. To achieve this, Calling Cards uses the piggyBac transposase to insert self-reporting transposon (SRT) "Calling Cards" into the genome, leaving permanent marks at interaction sites. Calling Cards can be deployed in a variety of in vitro and in vivo biological systems to study gene regulatory networks involved in development, aging, and disease. Out of the box, it assesses enhancer usage but can be adapted to profile specific transcription factor binding with custom transcription factor (TF)-piggyBac fusion proteins. The Calling Cards workflow has five main stages: delivery of Calling Card reagents, sample preparation, library preparation, sequencing, and data analysis. Here, we first present a comprehensive guide for experimental design, reagent selection, and optional customization of the platform to study additional TFs. Then, we provide an updated protocol for the five steps, using reagents that improve throughput and decrease costs, including an overview of a newly deployed computational pipeline. This protocol is designed for users with basic molecular biology experience to process samples into sequencing libraries in 1-2 days. Familiarity with bioinformatic analysis and command line tools is required to set up the pipeline in a high-performance computing environment and to conduct downstream analyses. Basic Protocol 1: Preparation and delivery of Calling Cards reagentsBasic Protocol 2: Sample preparationBasic Protocol 3: Sequencing library preparationBasic Protocol 4: Library pooling and sequencingBasic Protocol 5: Data analysis.
Calling cards technology using self-reporting transposons enables the identification of DNA-protein interactions through RNA sequencing. Although immensely powerful, current implementations of calling cards in bulk experiments on populations of cells are technically cumbersome and require many replicates to identify independent insertions into the same genomic locus. Here, we have drastically reduced the cost and labor requirements of calling card experiments in bulk populations of cells by introducing a DNA barcode into the calling card itself. An additional barcode incorporated during reverse transcription enables simultaneous transcriptome measurement in a facile and affordable protocol. We demonstrate that barcoded self-reporting transposons recover in vitro binding sites for four basic helix-loop-helix transcription factors with important roles in cell fate specification: ASCL1, MYOD1, NEUROD2, and NGN1. Further, simultaneous calling cards and transcriptional profiling during transcription factor overexpression identified both binding sites and gene expression changes for two of these factors. Lastly, we demonstrated barcoded calling cards can record binding in vivo in the mouse brain. In sum, RNA-based identification of transcription factor binding sites and gene expression through barcoded self-reporting transposon calling cards and transcriptomes is an efficient and powerful method to infer gene regulatory networks in a population of cells.
Summary Transposon calling cards is a genomic assay for identifying transcription factor binding sites in both bulk and single cell experiments. Here we describe the qBED format, an open, text-based standard for encoding and analyzing calling card data. In parallel, we introduce the qBED track on the WashU Epigenome Browser, a novel visualization that enables researchers to inspect calling card data in their genomic context. Finally, through examples, we demonstrate that qBED files can be used to visualize non-calling card datasets, such as CADD scores and GWAS/eQTL hits, and may have broad utility to the genomics community. Availability and Implementation The qBED track is available on the WashU Epigenome Browser ( http://epigenomegateway.wustl.edu/browser ), beginning with version 46. Source code for the WashU Epigenome Browser with qBED support is available on GitHub ( http://github.com/arnavm/eg-react and http://github.com/lidaof/eg-react ). We have also released a tutorial on how to upload qBED data to the browser ( dx.doi.org/10.17504/protocols.io.bca8ishw ).
Sex can be an important determinant of cancer phenotype, and exploring sex-biased tumor biology holds promise for identifying novel therapeutic targets and new approaches to cancer treatment. In an established isogenic murine model of glioblastoma (GBM), we discovered correlated transcriptome-wide sex differences in gene expression, H3K27ac marks, large Brd4-bound enhancer usage, and Brd4 localization to Myc and p53 genomic binding sites. These sex-biased gene expression patterns were also evident in human glioblastoma stem cells (GSCs). These observations led us to hypothesize that Brd4-bound enhancers might underlie sex differences in stem cell function and tumorigenicity in GBM. We found that male and female GBM cells exhibited sex-specific responses to pharmacological or genetic inhibition of Brd4. Brd4 knockdown or pharmacologic inhibition decreased male GBM cell clonogenicity and in vivo tumorigenesis while increasing both in female GBM cells. These results were validated in male and female patient-derived GBM cell lines. Furthermore, analysis of the Cancer Therapeutic Response Portal of human GBM samples segregated by sex revealed that male GBM cells are significantly more sensitive to BET (bromodomain and extraterminal) inhibitors than are female cells. Thus, Brd4 activity is revealed to drive sex differences in stem cell and tumorigenic phenotypes, which can be abrogated by sex-specific responses to BET inhibition. This has important implications for the clinical evaluation and use of BET inhibitors.
Transcription factors (TFs) enact precise regulation of gene expression through site-specific, genome-wide binding. Common methods for TF-occupancy profiling, such as chromatin immuno-precipitation, are limited by requirement of TF-specific antibodies and provide only end-point snapshots of TF binding. Alternatively, TF-tagging techniques, in which a TF is fused to a DNA-modifying enzyme that marks TF-binding events across the genome as they occur, do not require TF-specific antibodies and offer the potential for unique applications, such as recording of TF occupancy over time and cell type specificity through conditional expression of the TF-enzyme fusion. Here, we create a viral toolkit for one such method, calling cards, and demonstrate that these reagents can be delivered to the live mouse brain and used to report TF occupancy. Further, we establish a Cre-dependent calling cards system and, in proof-of-principle experiments, show utility in defining cell type-specific TF profiles and recording and integrating TF-binding events across time. This versatile approach will enable unique studies of TF-mediated gene regulation in live animal models.
This document explains how to visualize calling card insertions as wells as density and peak tracks on the WashU Epigenome Browser.
Cellular heterogeneity confounds in situ assays of transcription factor (TF) binding. Single-cell RNA sequencing (scRNA-seq) deconvolves cell types from gene expression, but no technology links cell identity to TF binding sites (TFBS) in those cell types. We present self-reporting transposons (SRTs) and use them in single-cell calling cards (scCC), a novel assay for simultaneously measuring gene expression and mapping TFBS in single cells. The genomic locations of SRTs are recovered from mRNA, and SRTs deposited by exogenous, TF-transposase fusions can be used to map TFBS. We then present scCC, which map SRTs from scRNA-seq libraries, simultaneously identifying cell types and TFBS in those same cells. We benchmark multiple TFs with this technique. Next, we use scCC to discover BRD4-mediated cell-state transitions in K562 cells. Finally, we map BRD4 binding sites in the mouse cortex at single-cell resolution, establishing a new method for studying TF biology in situ.
Sex can be an important determinant of cancer phenotype and exploring sex-specific tumor biology holds promise for identifying novel therapeutic targets and new approaches to cancer treatment. In an established model of glioblastoma, we discovered transcriptome-wide sexual dimorphism in gene-expression that was concordant with sex differences in H3K27ac marks, large Brd4-bound enhancer usage, and Brd4 co-localization with Myc and p53. The sex-specific enhancer usage drove sex differences in stem cell function and tumorigenicity. Moreover, male and female GBM cells exhibited opposing responses to pharmacological or genetic inhibition of Brd4. Brd4 knockdown or inhibition decreased male GBM cell clonogenicity and in vivo tumorigenesis, while increasing both in female GBM cells. These results were validated in male and female patient-derived GBM cell lines. Thus, for the first time, Brd4 activity is revealed to drive a sexually dimorphic stem cell and tumorigenic phenotype, resulting in diametrically opposite responses to BET inhibition in male and female glioblastoma cells. This has critical implications for the clinical evaluation and use of BET inhibitors. Citation Format: Najla Kfoury-Beaumont, Zongtai Qi, Michael Wilkinson, Lauren Broestl, Kristopher Berrett, Arnav Moudgil, Sumithra Sankararaman, Xuhua Chen, Jay Gertz, Robi Mitra, Joshua B. Rubin. Brd4-bound enhancers drive critical sex differences in glioblastoma [abstract]. In: Proceedings of the Annual Meeting of the American Association for Cancer Research 2020; 2020 Apr 27-28 and Jun 22-24. Philadelphia (PA): AACR; Cancer Res 2020;80(16 Suppl):Abstract nr 3658.
This protocol describes how to all peaks on mammalian calling card data using either undirected, or transcription factor fusions, to the piggyBac transposase. It is applicable for both bulk as well as single cell calling card data.
This protocol describes how to create calling card libraries from bulk RNA. This protocol assumes you have successfully transformed cells with piggyBac self-reporting transposons and either undirected piggyBac transposase or your favorite transcription factor (YFTF) fused to piggyBac. Your cells are now ready for RNA extraction, SRT amplification, and library preparation.
This protocol describes how to create calling card libraries from single cell RNA. We assume you have successfully transformed cells with piggyBac self-reporting transposons and either undirected piggyBac transposase or your favorite transcription factor (YFTF) fused to piggyBac. We also assume you have optimized the dissociation protocol for your specific cells or tissues and can generate single cell suspensions.
AbstractTranscription factors (TFs) play a central role in the regulation of gene expression, controlling everything from cell fate decisions to activity dependent gene expression. However, widely-used methods for TF profiling in vivo (e.g. ChIP-seq) yield only an aggregated picture of TF binding across all cell types present within the harvested tissue; thus, it is challenging or impossible to determine how the same TF might bind different portions of the genome in different cell types, or even to identify its binding events at all in rare cell types in a complex tissue such as the brain. Here we present a versatile methodology, FLEX Calling Cards, for the mapping of TF occupancy in specific cell types from heterogenous tissues. In this method, the TF of interest is fused to a hyperactive piggyBac transposase (hypPB), and this bipartite gene is delivered, along with donor transposons, to mouse tissue via a Cre-dependent adeno-associated virus (AAV). The fusion protein is expressed in Cre-expressing cells where it inserts transposon “Calling Cards” near to TF binding sites. These transposons permanently mark TF binding events and can be mapped using high-throughput sequencing. Alternatively, unfused hypPB interacts with and records the binding of the super enhancer (SE)-associated bromodomain protein, Brd4. To demonstrate the FLEX Calling Card method, we first show that donor transposon and transposase constructs can be efficiently delivered to the postnatal day 1 (P1) mouse brain with AAV and that insertion profiles report TF occupancy. Then, using a Cre-dependent hypPB virus, we show utility of this tool in defining cell type-specific TF profiles in multiple cell types of the brain. This approach will enable important cell type-specific studies of TF-mediated gene regulation in the brain and will provide valuable insights into brain development, homeostasis, and disease.
Here we present a computational pipeline for processing bulk RNA calling card data. These data will have been generated from transfection/-duction of either undirected piggyBac transposase or your favorite transcription factor (YFTF) fused to piggyBac. Multiple biological replicates should have been generated, each with a unique combination of primer barcode and index sequences. This workflow demonstrates how to analyze a single replicate; the workflow can be parallelized on distributed computing architectures (e.g. slurm).
Here we present a computational pipeline for processing single cell calling card (scCC) data. These data will have been generated from single cell RNA-seq libraries following transfection/-duction of either undirected piggyBac transposase or your favorite transcription factor (YFTF) fused to piggyBac. This workflow demonstrates how to process scCC sequencing data derived from a 10x Chromium-based scCC library; the workflow can be parallelized on distributed computing architectures (e.g. slurm).
Transposon calling cards can identify transcription factor (TF) binding sites. This involves fusing your favorite TF (YFTF) to the hyperactive piggyBac transposase (HyPBase). This is delivered to cells in conjunction with a piggyBac transposon. The TF will visit sites in the genome and YFTF-HyPBase will deposit transposons near binding sites. We then generate sequencing libraries to map the genome-wide localization of transposons. Finally, we identify significant clusters of insertions to identify TF binding sites.
It has long been debated whether natural selection acts primarily upon individual organisms, or whether it also commonly acts upon higher-level entities such as lineages. Two arguments against the effectiveness of long-term selection on lineages have been (i) that long-term evolutionary outcomes will not be sufficiently predictable to support a meaningful long-term fitness and (ii) that short-term selection on organisms will almost always overpower long-term selection. Here, we use a computational model of protein folding and binding called ‘lattice proteins’. We quantify the long-term evolutionary success of lineages with two metrics called thek-fitness andk-survivability. We show that long-term outcomes aresurprisingly predictablein this model: only a small fraction of the possible outcomes are ever realized in multiple replicates. Furthermore, the long-term fitness of a lineage depends only partly on its short-term fitness; other factors are also important, including the ‘evolvability’ of a lineage—its capacity to produce adaptive variation. In a system with a distinct short-term and long-term fitness, evolution need not be ‘short-sighted’: lineages may be selected for their long-term properties, sometimes in opposition to short-term selection. Similar evolutionary basins of attraction have been observedin vivo, suggesting that natural biological lineages will also have a predictive long-term fitness.
Human leucocyte antigen (HLA) genes play an important role in the success of organ transplantation and are associated with autoimmune and infectious diseases. Current DNA-based genotyping methods, including Sanger sequence-based typing (SSBT), have identified a high degree of polymorphism. This level of polymorphism makes high-resolution HLA genotyping challenging, resulting in ambiguous typing results due to an inability to resolve phase and/or defining polymorphisms lying outside the region amplified. Next-generation sequencing (NGS) may resolve the issue through the combination of clonal amplification, which provides phase information, and the ability to sequence larger regions of genes, including introns, without the additional effort or cost associated with current methods. The NGS HLA sequencing project of the 16IHIW aimed to discuss the different approaches to (i) template preparation including short- and long-range PCR amplicons, exome capture and whole genome; (ii) sequencing platforms, including GS 454 FLX, Ion Torrent PGM, Illumina MiSeq/HiSeq and Pacific Biosciences SMRT; (iii) data analysis, specifically allele-calling software. The pilot studies presented at the workshop demonstrated that although individual sequencers have very different performance characteristics, all produced sequence data suitable for the resolution of HLA genotyping ambiguities. The developments presented at this workshop clearly highlight the potential benefits of NGS in the HLA laboratory.