Virtual screening asks which molecules, among an enormous space of drug-like chemistry, are worth synthesizing and testing against a protein target. Most modern methods answer by building and scoring an explicit three-dimensional pose through molecular docking, or the co-folding models that now approach experimental accuracy. Building these poses presumes a well-defined pocket, and the non-orthosteric, cryptic, and intrinsically disordered sites where unexplored ligandability lies offer none. There, these methods fail to generalize. Here we present Ptarmigan-1, a contrastive model that co-embeds the residues of a protein with candidate small molecules in a shared latent space, from sequence and two-dimensional chemistry alone, and without ever constructing a pose. Engagement reduces to the proximity of precomputed embeddings. Freed from the pose, Ptarmigan-1 trains directly on chemoproteomic and bioactivity data of mixed resolution, scores a compound in ten milliseconds rather than the tens of seconds a co-folding model demands, and resolves each prediction to the residues a compound engages. On well-folded, orthosteric targets it performs comparably to a collection of co-folding and docking models, and on covalent, cryptic, and disordered sites it matches or exceeds them. It localizes reversible and covalent inhibitors to the pockets they engage, even for targets withheld from training, and screens the entire human proteome against a library of 3.4 billion compounds in under a day. By decoupling molecular recognition from structure, Ptarmigan-1 recasts virtual screening as a reusable index that continuously improves as data accumulate.
Supplementary Figure S1. Generation of LSL-EZH2 conditional transgenic mice. Supplementary Figure S2. H3K27ac super enhancer analysis in EZH2 mouse lung tumors. Supplementary Figure S3. Human NSCLC cells with low levels of EZH2 are not sensitive to disruption of EZH2. Supplementary Figure S4. Development and characterization of JQEZ5. Supplementary Figure S5. Cellular and in vivo characterization of JQEZ5. Supplementary Table S1. Summary of LSL-EZH2 mouse models. Supplementary Table S2. Summary of lung adenocarcinoma in LSL-EZH2 mouse models.
Advances in library-based methods for peptide detection from data-independent acquisition (DIA) mass spectrometry have made it possible to detect and quantify tens of thousands of peptides in a single mass spectrometry run. However, many of these methods rely on a comprehensive, high-quality spectral library containing information about the expected retention time and fragmentation patterns of peptides in the sample. Empirical spectral libraries are often generated through data-dependent acquisition and may suffer from biases as a result. Spectral libraries can be generated in silico, but these models are not trained to handle all possible post-translational modifications. Here, we propose a false discovery rate-controlled spectrum-centric search workflow to generate spectral libraries directly from gas-phase fractionated DIA tandem mass spectrometry data. We demonstrate that this strategy is able to detect phosphorylated peptides and can be used to generate a spectral library for accurate peptide detection and quantitation in wide-window DIA data. We compare the results of this search workflow to other library-free approaches and demonstrate that our search is competitive in terms of accuracy and sensitivity. These results demonstrate that the proposed workflow has the capacity to generate spectral libraries while avoiding the limitations of other methods.
Advances in library-based methods for peptide detection from data independent acquisition (DIA) mass spectrometry have made it possible to detect and quantify tens of thousands of peptides in a single mass spectrometry run. However, many of these methods rely on a comprehensive, high quality spectral library containing information about the expected retention time and fragmentation patterns of peptides in the sample. Empirical spectral libraries are often generated through data-dependent acquisition and may suffer from biases as a result. Spectral libraries can be generated in silico but these models are not trained to handle all possible post-translational modifications. Here, we propose a false discovery rate controlled spectrum-centric search workflow to generate spectral libraries directly from gas-phase fractionated DIA tandem mass spectrometry data. We demonstrate that this strategy is able to detect phosphorylated peptides and can be used to generate a spectral library for accurate peptide detection and quantitation in wide window DIA data. We compare the results of this search workflow to other library-free approaches and demonstrate that our search is competitive in terms of accuracy and sensitivity. These results demonstrate that the proposed workflow has the capacity to generate spectral libraries while avoiding the limitations of other methods.
Newly synthesized histone H4 that is incorporated into chromatin during DNA replication is acetylated on lysines 5 and 12. Histone deacetylase 1 (HDAC1) and HDAC2 are responsible for reducing H4 acetylation as chromatin matures. Using CRISPR-Cas9-generated hdac1- or hdac2-null fibroblasts, we determined that HDAC1 and HDAC2 do not fully compensate for each other in removing de novo acetyls on H4 in vivo. Proteomics of nascent chromatin and proximity ligation assays with newly replicated DNA revealed the binding of ATAD2, a bromodomain-containing posttranslational modification (PTM) reader that recognizes acetylated H4. ATAD2 is a transcription facilitator overexpressed in several cancers and in the simian virus 40 (SV40)-transformed human fibroblast model cell line used in this study. The recruitment of ATAD2 to nascent chromatin was increased in hdac2 cells over the wild type, and ATAD2 depletion reduced the levels of nascent chromatin-associated, acetylated H4 in wild-type and hdac2 cells. We propose that overexpressed ATAD2 shifts the balance of H4 acetylation by protecting this mark from removal and that HDAC2 but not HDAC1 can effectively compete with ATAD2 for the target acetyls. ATAD2 depletion also reduced global RNA synthesis and nascent DNA-associated RNA. A moderate dependence on ATAD2 for replication fork progression was noted only for hdac2 cells overexpressing the protein.
Transcription factors and other chromatin-associated proteins are difficult to quantify comprehensively. Here, we combine facile nuclear sub-fractionation with data-independent acquisition mass spectrometry to achieve rapid, sensitive, and highly parallel quantification of the nuclear proteome in human cells. We apply this approach to quantify the response to acute degradation of BET bromodomains, revealing unexpected chromatin regulatory dynamics. The method is simple and enables system-level study of previously inaccessible chromatin and genome regulators.
Non-invasive epigenome editing is a promising strategy for engineering gene expression programs, yet potency, specificity, and persistence remain challenging. Here we show that effective epigenome editing is gated at single-base precision via 'keyhole' sites in endogenous regulatory DNA. Synthetic repressors targeting promoter keyholes can ablate gene expression in up to 99% of primary cells with single-gene specificity and can seamlessly repress multiple genes in combination. Transient exposure of primary T cells to keyhole repressors confers mitotically heritable silencing that persists to the limit of primary cultures in vitro and for at least 4 weeks in vivo, enabling manufacturing of cell products with enhanced therapeutic efficacy. DNA recognition and effector domains can be encoded as separate proteins that reassemble at keyhole sites and function with the same efficiency as single chain effectors, enabling gated control and rapid screening for novel functional domains that modulate endogenous gene expression patterns. Our results provide a powerful and exponentially flexible system for programming gene expression and therapeutic cell products.
Currently, there is a growing need for culturing hematopoietic stem/progenitor cells (HSPCs) in vitro for various clinical applications including gene therapy. Compared with cord blood (CB) CD34+ HSPCs, it is more challenging to maintain or expand CD34+ peripheral blood mobilized stem/progenitor cells (PBSCs) ex vivo. To fill this knowledge gap, we have systematically surveyed 466 small-molecule drug compounds for their potential in cytokine-dependent expansion of human CD34+CD90+ HSPCs. We found that epigenetic modifiers, especially histone deacetylase inhibitors (HDACis), could preferentially maintain and expand these cells. In particular, treatment of CD34+ PBSCs with a single dose of HDACi trichostatin A (TSA) at a concentration of 50 nmol/L ex vivo yielded the greatest expansion (11.7-fold) of CD34+CD90+ cells when compared with the control (dimethyl sulfoxide [DMSO] plus cytokines) group. Additionally, TSA-treated PBSC CD34+ cells had a statistically significant higher engraftment rate than the control-treated group in xenotransplantation experiments. Mechanistically, TSA treatment was associated with increased expression of HSPC-related genes such as GATA2 and SALL4. Furthermore, TSA-mediated CD34+CD90+ expansion was reduced by downregulation of SALL4 but not GATA2. Overall, we have developed a robust, short-term (5-day), PBSC ex vivo maintenance/expansion culture technique and found that the HDACi-TSA/SALL4 axis is important for the biological process.
Sequencing-based technologies cannot measure post-transcriptional dynamics of the nuclear proteome, but unbiased mass-spectrometry measurements of nuclear proteins remain difficult. In this work, we have combined facile nuclear sub-fractionation approaches with data-independent acquisition mass spectrometry to improve detection and quantification of nuclear proteins in human cells and tissues. Nuclei are isolated and subjected to a series of extraction conditions that enrich for nucleoplasm, euchromatin, heterochromatin and nuclear-membrane associated proteins. Using this approach, we can measure peptides from over 70% of the expressed nuclear proteome. As we are physically separating chromatin compartments prior to analysis, proteins can be assigned into functional chromatin environments that illuminate systems-wide nuclear protein dynamics. To validate the integrity of nuclear sub-compartments, immunofluorescence confirms the presence of key markers during chromatin extraction. We then apply this method to study the nuclear proteome-wide response to during pharmacological degradation of the BET bromodomain proteins. BET degradation leads to widespread changes in chromatin composition, and we discover global HDAC1/2-mediated remodeling of chromatin previously bound by BET bromodomains. In summary, we have developed a technology for reproducible, comprehensive characterization of the nuclear proteome to observe the systems-wide nuclear protein dynamics.
Developmental transitions are guided by master regulatory transcription factors. During adipogenesis, a transcriptional cascade culminates in the expression of PPARγ and C/EBPα, which orchestrate activation of the adipocyte gene expression program. However, the coactivators controlling PPARγ and C/EBPα expression are less well characterized. Here, we show the bromodomain-containing protein, BRD4, regulates transcription of PPARγ and C/EBPα. Analysis of BRD4 chromatin occupancy reveals that induction of adipogenesis in 3T3L1 fibroblasts provokes dynamic redistribution of BRD4 to de novo super-enhancers proximal to genes controlling adipocyte differentiation. Inhibition of the bromodomain and extraterminal domain (BET) family of bromodomain-containing proteins impedes BRD4 occupancy at these de novo enhancers and disrupts transcription of Pparg and Cebpa, thereby blocking adipogenesis. Furthermore, silencing of these BRD4-occupied distal regulatory elements at the Pparg locus by CRISPRi demonstrates a critical role for these enhancers in the control of Pparg gene expression and adipogenesis in 3T3L1s. Together, these data establish BET bromodomain proteins as time- and context-dependent coactivators of the adipocyte cell state transition.
Adipocyte turnover in adulthood is low, suggesting that the cellular source of new adipocytes, the adipocyte progenitor (AP), resides in a state of relative quiescence. Yet the core transcriptional regulatory circuitry (CRC) responsible for establishing a quiescent state and the physiological significance of AP quiescence are incompletely understood. Here, we integrate transcriptomic data with maps of accessible chromatin in primary APs, implicating the orphan nuclear receptor NR4A1 in AP cell-state regulation. NR4A1 gain and loss of function in APs ex vivo decreased and enhanced adipogenesis, respectively. Adipose tissue of Nr4a1(-/-) mice demonstrated higher proliferative and adipogenic capacity compared with that of WT mice. Transplantation of Nr4a1(-/-) APs into the subcutaneous adipose tissue of WT obese recipients improved metrics of glucose homeostasis relative to administration of WT APs. Collectively, these data identify NR4A1 as a previously unrecognized constitutive regulator of AP quiescence and suggest that augmentation of adipose tissue plasticity may attenuate negative metabolic sequelae of obesity.
Enhancer profiling is a powerful approach for discovering cis-regulatory elements that define the core transcriptional regulatory circuits of normal and malignant cells. Gene control through enhancer activity is often dominated by a subset of lineage-specific transcription factors. By integrating measures of chromatin accessibility and enrichment for H3K27 acetylation, we have generated regulatory landscapes of chronic lymphocytic leukemia (CLL) samples and representative cell lines. With super enhancer-based modeling of regulatory circuits and assessments of transcription factor dependencies, we discover that the essential super enhancer factor PAX5 dominates CLL regulatory nodes and is essential for CLL cell survival. Targeting enhancer signaling via BET bromodomain inhibition disrupts super enhancer-dependent gene expression with selective effects on CLL core regulatory circuitry, conferring potent anti-tumor activity.
Regulation of gene expression through binding of transcription factors (TFs) to cis-regulatory elements is highly complex in mammalian cells. Genome-wide measurement technologies provide new means to understand this regulation, and models of TF regulatory networks have been built with the goal of identifying critical factors. Here, we report a network model of transcriptional regulation between TFs constructed by integrating genomewide identification of active enhancers and regions of focal DNA accessibility. Network topology is confirmed by published TF ChIP-seq data. By considering multiple methods of TF prioritization following network construction, we identify master TFs in well-studied cell types, and these networks provide better prioritization than networks only considering promoter-proximal accessibility peaks. Comparisons between networks from similar cell types show stable connectivity of most TFs, while master regulator TFs show dramatic changes in connectivity and centrality. Applying this method to study chronic lymphocytic leukemia, we prioritized several network TFs amenable to pharmacological perturbation and show that compounds targeting these TFs show comparable efficacy in CLL cell lines to FDA-approved therapies. The construction of transcriptional regulatory network (TRN) models can predict the interactions between individual TFs and predict critical TFs for development or disease.
Merkel cell carcinoma (MCC) frequently contains integrated copies of Merkel cell polyomavirus DNA that express a truncated form of Large T antigen (LT) and an intact Small T antigen (ST). While LT binds RB and inactivates its tumor suppressor function, it is less clear how ST contributes to MCC tumorigenesis. Here we show that ST binds specifically to the MYC homolog MYCL (L-MYC) and recruits it to the 15-component EP400 histone acetyltransferase and chromatin remodeling complex. We performed a large-scale immunoprecipitation for ST and identified co-precipitating proteins by mass spectrometry. In addition to protein phosphatase 2A (PP2A) subunits, we identified MYCL and its heterodimeric partner MAX plus the EP400 complex. Immunoprecipitation for MAX and EP400 complex components confirmed their association with ST. We determined that the ST-MYCL-EP400 complex binds together to specific gene promoters and activates their expression by integrating chromatin immunoprecipitation with sequencing (ChIP-seq) and RNA-seq. MYCL and EP400 were required for maintenance of cell viability and cooperated with ST to promote gene expression in MCC cell lines. A genome-wide CRISPR-Cas9 screen confirmed the requirement for MYCL and EP400 in MCPyV-positive MCC cell lines. We demonstrate that ST can activate gene expression in a EP400 and MYCL dependent manner and this activity contributes to cellular transformation and generation of induced pluripotent stem cells.
Genomic sequencing has driven precision-based oncology therapy; however, the genetic drivers of many malignancies remain unknown or non-targetable, so alternative approaches to the identification of therapeutic leads are necessary. Ependymomas are chemotherapy-resistant brain tumours, which, despite genomic sequencing, lack effective molecular targets. Intracranial ependymomas are segregated on the basis of anatomical location (supratentorial region or posterior fossa) and further divided into distinct molecular subgroups that reflect differences in the age of onset, gender predominance and response to therapy. The most common and aggressive subgroup, posterior fossa ependymoma group A (PF-EPN-A), occurs in young children and appears to lack recurrent somatic mutations. Conversely, posterior fossa ependymoma group B (PF-EPN-B) tumours display frequent large-scale copy number gains and losses but have favourable clinical outcomes. More than 70% of supratentorial ependymomas are defined by highly recurrent gene fusions in the NF-κB subunit gene RELA (ST-EPN-RELA), and a smaller number involve fusion of the gene encoding the transcriptional activator YAP1 (ST-EPN-YAP1). Subependymomas, a distinct histologic variant, can also be found within the supratetorial and posterior fossa compartments, and account for the majority of tumours in the molecular subgroups ST-EPN-SE and PF-EPN-SE. Here we describe mapping of active chromatin landscapes in 42 primary ependymomas in two non-overlapping primary ependymoma cohorts, with the goal of identifying essential super-enhancer-associated genes on which tumour cells depend. Enhancer regions revealed putative oncogenes, molecular targets and pathways; inhibition of these targets with small molecule inhibitors or short hairpin RNA diminished the proliferation of patient-derived neurospheres and increased survival in mouse models of ependymomas. Through profiling of transcriptional enhancers, our study provides a framework for target and drug discovery in other cancers that lack known genetic drivers and are therefore difficult to treat.
Medulloblastoma is a highly malignant paediatric brain tumour, often inflicting devastating consequences on the developing child. Genomic studies have revealed four distinct molecular subgroups with divergent biology and clinical behaviour. An understanding of the regulatory circuitry governing the transcriptional landscapes of medulloblastoma subgroups, and how this relates to their respective developmental origins, is lacking. Here, using H3K27ac and BRD4 chromatin immunoprecipitation followed by sequencing (ChIP-seq) coupled with tissue-matched DNA methylation and transcriptome data, we describe the active cis-regulatory landscape across 28 primary medulloblastoma specimens. Analysis of differentially regulated enhancers and super-enhancers reinforced inter-subgroup heterogeneity and revealed novel, clinically relevant insights into medulloblastoma biology. Computational reconstruction of core regulatory circuitry identified a master set of transcription factors, validated by ChIP-seq, that is responsible for subgroup divergence, and implicates candidate cells of origin for Group 4. Our integrated analysis of enhancer elements in a large series of primary tumour samples reveals insights into cis-regulatory architecture, unrecognized dependencies, and cellular origins.
Abstract As a master regulator of chromatin function, the lysine methyltransferase EZH2 orchestrates transcriptional silencing of developmental gene networks. Overexpression of EZH2 is commonly observed in human epithelial cancers, such as non–small cell lung carcinoma (NSCLC), yet definitive demonstration of malignant transformation by deregulated EZH2 remains elusive. Here, we demonstrate the causal role of EZH2 overexpression in NSCLC with new genetically engineered mouse models of lung adenocarcinoma. Deregulated EZH2 silences normal developmental pathways, leading to epigenetic transformation independent of canonical growth factor pathway activation. As such, tumors feature a transcriptional program distinct from KRAS- and EGFR-mutant mouse lung cancers, but shared with human lung adenocarcinomas exhibiting high EZH2 expression. To target EZH2-dependent cancers, we developed a potent open-source EZH2 inhibitor, JQEZ5, that promoted the regression of EZH2-driven tumors in vivo, confirming oncogenic addiction to EZH2 in established tumors and providing the rationale for epigenetic therapy in a subset of lung cancer. Significance: EZH2 overexpression induces murine lung cancers that are similar to human NSCLC with high EZH2 expression and low levels of phosphorylated AKT and ERK, implicating biomarkers for EZH2 inhibitor sensitivity. Our EZH2 inhibitor, JQEZ5, promotes regression of these tumors, revealing a potential role for anti-EZH2 therapy in lung cancer. Cancer Discov; 6(9); 1006–21. ©2016 AACR. See related commentary by Frankel et al., p. 949. This article is highlighted in the In This Issue feature, p. 932
A small set of core transcription factors (TFs) dominates control of the gene expression program in embryonic stem cells and other well-studied cellular models. These core TFs collectively regulate their own gene expression, thus forming an interconnected auto-regulatory loop that can be considered the core transcriptional regulatory circuitry (CRC) for that cell type. There is limited knowledge of core TFs, and thus models of core regulatory circuitry, for most cell types. We recently discovered that genes encoding known core TFs forming CRCs are driven by super-enhancers, which provides an opportunity to systematically predict CRCs in poorly studied cell types through super-enhancer mapping. Here, we use super-enhancer maps to generate CRC models for 75 human cell and tissue types. These core circuitry models should prove valuable for further investigating cell-type–specific transcriptional regulation in healthy and diseased cells.