Abstract The functional impact of a large portion of the human genome known as “dark matter DNA”, which is composed mainly of repeat sequences, remains unknown. The genome also encodes many putative and poorly characterized transcription factors. Here, we determine genomic binding locations of 166 poorly characterized human transcription factors in living cells. Nearly half of them associate strongly with known regulatory regions such as promoters and enhancers, frequently co-localizing with each other at conserved motif matches. The other half often associate with genomic dark matter, however, at largely non-overlapping (i.e., unique) sites, via intrinsic sequence recognition. Fifty-four of the latter half, which we term dark transcription factors , mainly bind within regions of closed chromatin, with each recognizing a unique set of repeat sequences. The dark transcription factors include many KZNFs, which are known to bind and silence transposable elements, and other transcription factors with apparent repressive functions. Others may be pioneer transcription factors. For example, we find that induction of TPRX1, a known regulator of zygotic preimplantation, leads to chromatin opening at many of its binding sites in the dark matter genome.
Gene expression is regulated by transcription factors (TFs), which recognize specific DNA sequence motifs. Several hundred putative human TFs, identified mainly by an apparent DNA-binding domain, lack known binding motifs1. Furthermore, even for well-characterized TFs, it remains controversial the degree to which motifs accurately reflect binding sites in living cells2. Here we describe a systematic effort ('Codebook') to determine the sequence specificity of 332 putative and poorly characterized human TFs. More than 4,000 independent experiments, encompassing multiple in vitro and in vivo assays, produced motifs for just over half (177; 53%) of the TFs, of which most are associated with only a single protein. These results extend the vocabulary of sequence recognition encoded by human TFs by around 130 distinct motifs. Moreover, binding motifs identified in vitro are strongly enriched in cellular binding sites. Collectively, the data reveal tens of thousands of previously unknown, conserved and direct TF-binding sites across the human genome. These sites are concentrated in promoter regions and are predictive of gene expression. In summary, this new codebook provides an important step forward in decoding the human genome.
Abstract Motivation DNA motifs recognised by transcription factors are typically represented as position weight matrices (PWMs), assuming independent contributions of individual nucleotides to protein binding specificity. Many alternative models accounting for correlations of positional contributions have been introduced in the past decades. However, performance gains have generally not outweighed the advantages of simplicity, interpretability, and practical applicability of PWMs with the well-established codebase. Existing software tools and motif databases provide multiple non-identical PWMs for the same transcription factor or even for the same dataset. It remains a practical question whether these PWMs can be effectively combined into a single improved model. Results Here we describe ArChIPelago ( https://github.com/autosome-ru/ArChIPelago ), a computational framework that combines multiple PWMs into a joint model using classic machine learning techniques, from linear regression to ensembles of decision trees. We show that such a combination improves prediction of transcription factor binding sites in genomic sequences. With a diverse collection of 704 ChIP-Seq datasets spanning 36 orthologous human and mouse transcription factors of diverse structural families, we show that ArChIPelago consistently outperforms the best available individual mono- and dinucleotide PWMs as well as sparse local inhomogeneous mixture models. Furthermore, using both human and mouse data, we demonstrate that PWM ensembles are capable of making reliable cross-species predictions.
Human papillomavirus-positive (HPV+) head and neck squamous cell carcinoma (HNSCC) is a growing subset of cancer cases distinct from HPV-negative by fewer genetic mutations and prevalent epigenetic dysregulation. We mapped H3K27ac-marked super-enhancers (SEs) via ChIP-seq in HPV+ patient-derived xenografts (PDXs) and normal oropharyngeal mucosa, identifying tumor-specific SE domains (T-SEDs) enriched for transcription factors (TFs) including TP63, FOSL1, and JUND. These SE-associated TFs regulate key oncogenic pathways and are downregulated by BRD4 inhibition with JQ1, highlighting sensitivity to epigenetic modulation. RNA-seq data revealed coordinated dysregulation of enhancer RNAs and mRNAs near T-SEDs, linked to upregulated pathways including epithelial-mesenchymal transition and E2F targets. JQ1 treatment significantly repressed these tumor-specific pathways, suggesting a therapeutic potential for targeting SE-driven transcription in HPV+ HNSCC. This study underscores the critical role of SEs in epigenetic and transcriptional dysregulation in HPV+ HNSCC, revealing therapeutic targets and providing a framework for future mechanistic studies in this area.
A sequence motif representing the DNA-binding specificity of a transcription factor (TF) is commonly modelled with a positional weight matrix (PWM). Focusing on understudied human TFs, we processed results of 4,237 experiments for 394 TFs, assayed using five different experimental platforms. By human curation, we approved a subset of experiments that yielded consistent motifs across platforms and replicates, and evaluated quantitatively the cross-platform performance of PWMs obtained with ten motif discovery tools. Notably, nucleotide composition and information content are not correlated with motif performance and do not help in detecting underperformers, while motifs with low information content, in many cases, describe well the binding specificity assessed across different experimental platforms. By combining multiple PMWs into a random forest, we demonstrate the potential of accounting for multiple modes of TF binding. Finally, we present the Codebook Motif Explorer (https://mex.autosome.org), cataloguing motifs, benchmarking results, and the underlying experimental data.
DNA motif discovery and, particularly, computational modeling of transcription factor binding motifs, has been a mecca of algorithmic bioinformatics for several decades. Here, we report the results of the largest open community challenge in Inferring BInding Specificities (IBIS), where participants all over the world were invited to construct binding specificity models from multi-assay experimental data for poorly studied human transcription factors. The submissions were rigorously tested against a rich held-out dataset. Benchmarking demonstrated a consistent advantage of properly designed deep learning models over traditional positional weight matrices and other machine learning methods. Yet, the positional weight matrices displayed a surprisingly strong performance out of the box, being only slightly behind the best deep learning models. A post-challenge assessment of a selection of other deep learning methods further solidified this finding. IBIS highlights the power of benchmarking in finding adequate DNA motif representations, emphasizes the pros and cons of various machine learning methods applied to DNA motif modeling, and establishes a rich dataset, benchmarking protocols, and computational framework for a fair cross-platform evaluation of future models of transcription factor binding motifs in DNA sequences.
Most of the human genome is thought to be non-functional, and includes large segments often referred to as "dark matter" DNA. The genome also encodes hundreds of putative and poorly characterized transcription factors (TFs). We determined genomic binding locations of 166 uncharacterized human TFs in living cells. Nearly half of them associated strongly with known regulatory regions such as promoters and enhancers, often at conserved motif matches and co-localizing with each other. Surprisingly, the other half often associated with genomic dark matter, at largely unique sites, via intrinsic sequence recognition. Dozens of these, which we term "Dark TFs", mainly bind within regions of closed chromatin. Dark TF binding sites are enriched for transposable elements, and are rarely under purifying selection. Some Dark TFs are KZNFs, which contain the repressive KRAB domain, but many are not: the Dark TFs also include known or potential pioneer TFs. Compiled literature information supports that the Dark TFs exert diverse functions ranging from early development to tumor suppression. Thus, our results sheds light on a large fraction of previously uncharacterized human TFs and their unappreciated activities within the dark matter genome.
We describe an effort ("Codebook") to determine the sequence specificity of 332 putative and largely uncharacterized human transcription factors (TFs), as well as 61 control TFs. Nearly 5,000 independent experiments across multiple in vitro and in vivo assays produced motifs for just over half of the putative TFs analyzed (177, or 53%), of which most are unique to a single TF. The data highlight the extensive contribution of transposable elements to TF evolution, both in cis and trans, and identify tens of thousands of conserved, base-level binding sites in the human genome. The use of multiple assays provides an unprecedented opportunity to benchmark and analyze TF sequence specificity, function, and evolution, as further explored in accompanying manuscripts. 1,421 human TFs are now associated with a DNA binding motif. Extrapolation from the Codebook benchmarking, however, suggests that many of the currently known binding motifs for well-studied TFs may inaccurately describe the TF's true sequence preferences.
A DNA sequence pattern, or "motif", is an essential representation of DNA-binding specificity of a transcription factor (TF). Any particular motif model has potential flaws due to shortcomings of the underlying experimental data and computational motif discovery algorithm. As a part of the Codebook/GRECO-BIT initiative, here we evaluated at large scale the cross-platform recognition performance of positional weight matrices (PWMs), which remain popular motif models in many practical applications. We applied ten different DNA motif discovery tools to generate PWMs from the "Codebook" data comprised of 4,237 experiments from five different platforms profiling the DNA-binding specificity of 394 human proteins, focusing on understudied transcription factors of different structural families. For many of the proteins, there was no prior knowledge of a genuine motif. By benchmarking-supported human curation, we constructed an approved subset of experiments comprising about 30% of all experiments and 50% of tested TFs which displayed consistent motifs across platforms and replicates. We present the Codebook Motif Explorer (https://mex.autosome.org), a detailed online catalog of DNA motifs, including the top-ranked PWMs, and the underlying source and benchmarking data. We demonstrate that in the case of high-quality experimental data, most of the popular motif discovery tools detect valid motifs and generate PWMs, which perform well both on genomic and synthetic data. Yet, for each of the algorithms, there were problematic combinations of proteins and platforms, and the basic motif properties such as nucleotide composition and information content offered little help in detecting such pitfalls. By combining multiple PMWs in decision trees, we demonstrate how our setup can be readily adapted to train and test binding specificity models more complex than PWMs. Overall, our study provides a rich motif catalog as a solid baseline for advanced models and highlights the power of the multi-platform multi-tool approach for reliable mapping of DNA binding specificities.
We present a major update of the HOCOMOCO collection that provides DNA binding specificity patterns of 949 human transcription factors and 720 mouse orthologs. To make this release, we performed motif discovery in peak sets that originated from 14 183 ChIP-Seq experiments and reads from 2554 HT-SELEX experiments yielding more than 400 thousand candidate motifs. The candidate motifs were annotated according to their similarity to known motifs and the hierarchy of DNA-binding domains of the respective transcription factors. Next, the motifs underwent human expert curation to stratify distinct motif subtypes and remove non-informative patterns and common artifacts. Finally, the curated subset of 100 thousand motifs was supplied to the automated benchmarking to select the best-performing motifs for each transcription factor. The resulting HOCOMOCO v12 core collection contains 1443 verified position weight matrices, including distinct subtypes of DNA binding motifs for particular transcription factors. In addition to the core collection, HOCOMOCO v12 provides motif sets optimized for the recognition of binding sites in vivo and in vitro, and for annotation of regulatory sequence variants. HOCOMOCO is available at https://hocomoco12.autosome.org and https://hocomoco.autosome.org.
*** SUMMARY *** This dataset contains supplementary data accompanying HOCOMOCO v12 collection of DNA binding motifs for human and mouse transcription factors, https://hocomoco.autosome.org The contents include: - the complete initial set of motifs discovered from ChIP-Seq and HT-SELEX data; - the curated subset of motifs associated with distinct motif subtypes that were used in benchmarking; - the benchmarking results and the resulting final motif collections, including motif logos; - accompanying metadata. Please refer to the README and the HOCOMOCO website for further details.
We present an update of EpiFactors, a manually curated database providing information about epigenetic regulators, their complexes, targets, and products which is openly accessible at http://epifactors.autosome.org. An updated version of the EpiFactors contains information on 902 proteins, including 101 histones and protamines, and, as a main update, a newly curated collection of 124 lncRNAs involved in epigenetic regulation. The amount of publications concerning the role of lncRNA in epigenetics is rapidly growing. Yet, the resource that compiles, integrates, organizes, and presents curated information on lncRNAs in epigenetics is missing. EpiFactors fills this gap and provides data on epigenetic regulators in an accessible and user-friendly form. For 820 of the genes in EpiFactors, we include expression estimates across multiple cell types assessed by CAGE-Seq in the FANTOM5 project. In addition, the updated EpiFactors contains information on 73 protein complexes involved in epigenetic regulation. Our resource is practical for a wide range of users, including biologists, bioinformaticians and molecular/systems biologists.
We present ANANASTRA, https://ananastra.autosome.org, a web server for the identification and annotation of regulatory single-nucleotide polymorphisms (SNPs) with allele-specific binding events. ANANASTRA accepts a list of dbSNP IDs or a VCF file and reports allele-specific binding (ASB) sites of particular transcription factors or in specific cell types, highlighting those with ASBs significantly enriched at SNPs in the query list. ANANASTRA is built on top of a systematic analysis of allelic imbalance in ChIP-Seq experiments and performs the ASB enrichment test against background sets of SNPs found in the same source experiments as ASB sites but not displaying significant allelic imbalance. We illustrate ANANASTRA usage with selected case studies and expect that ANANASTRA will help to conduct the follow-up of GWAS in terms of establishing functional hypotheses and designing experimental verification.
Abstract The subset of head and neck squamous cell carcinomas (HNSCC) that are driven by infection with the Human Papilloma Virus (HPV+ HNSCC) are increasing in prevalence and are defined by distinct biological signatures and mechanisms, such as abnormal transcriptional regulation. Here, we use optimized protocols to perform ChIP-seq and RNA-seq on patient surgical samples, which we use to define promoter, typical-enhancer, and super-enhancer (SE) regulatory regions with differential chromatin states (i.e. H3K27ac signal) between tumor and normal samples. We then leverage publicly available transcription factor (TF) cistrome data to test 372 TFs for tumor or normal specific enrichment across these differential regulatory regions. We identify 31 tumor-specific TF enrichments, such as TP53/63/73, NKX2-1, JUNB/D, FOSL1/2, E2F1/3, KLF3/4/5, SMAD3/4, TEAD1/4, TFAP2A/C, HIF1A, PPARG, GRHL2, EPAS1, and SNAI2. A number of these enriched TFs, including TP63, TP73, FOSL1, and E2F1, show both significantly increased expression in tumors and significant transcriptional reductions after JQ1 treatment. These results provide novel insight into the regulatory biology of HPV+ HNSCC, and suggest that the transcriptional networks and feedback loops governing TF expression and dysregulation in HPV+ HNSCC can be targeted using JQ1 treatment. Citation Format: Fernando Zamuner, Spencer S. Chan, Michael Kessler, Ilya Vorontsov, Ludmila Danilova, Rossin Erbe, Dylan Kelley, Theresa Guo, Eddie Imada, Elana J. Fertig, Ivan Kulakovskiy, Alexander Favorov, Daria A. Gaykalova. Identification of transcription factor enrichments in HPV+ head and neck cancer using ChIP-seq and cistrome analysis [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2022; 2022 Apr 8-13. Philadelphia (PA): AACR; Cancer Res 2022;82(12_Suppl):Abstract nr 1459.
Somatic mutations in regulatory sites of human stem cells affect cell identity or cause malignant transformation. By mining the human genome for co-occurrence of mutations and transcription factor binding sites, we show that C/EBP binding sites are strongly enriched with [C > T]G mutations in cancer and adult stem cells, which is of special interest because C/EBPs regulate cell fate and differentiation. In vitro protein-DNA binding assay and structural modeling of the CEBPB-DNA complex show that the G center dot T mismatch in the core CG dinucleotide strongly enhances affinity of the binding site. We conclude that enhanced binding of C/EBPs shields CpG center dot TpG mismatches from DNA repair, leading to selective accumulation of [C > T]G mutations and consequent deterioration of the binding sites. This mechanism of targeted mutagenesis highlights the effect of a mutational process on certain regulatory sites and reveals the molecular basis of putative regulatory alterations in stem cells.
Sequence variants in gene regulatory regions alter gene expression and contribute to phenotypes of individual cells and the whole organism, including disease susceptibility and progression. Single-nucleotide variants in enhancers or promoters may affect gene transcription by altering transcription factor binding sites. Differential transcription factor binding in heterozygous genomic loci provides a natural source of information on such regulatory variants. We present a novel approach to call the allele-specific transcription factor binding events at single-nucleotide variants in ChIP-Seq data, taking into account the joint contribution of aneuploidy and local copy number variation, that is estimated directly from variant calls. We have conducted a meta-analysis of more than 7 thousand ChIP-Seq experiments and assembled the database of allele-specific binding events listing more than half a million entries at nearly 270 thousand single-nucleotide polymorphisms for several hundred human transcription factors and cell types. These polymorphisms are enriched for associations with phenotypes of medical relevance and often overlap eQTLs, making candidates for causality by linking variants with molecular mechanisms. Specifically, there is a special class of switching sites, where different transcription factors preferably bind alternative alleles, thus revealing allele-specific rewiring of molecular circuitry.
During translation, the rate of ribosome movement along mRNA varies. This leads to a non-uniform ribosome distribution along the transcript, depending on local mRNA sequence, structure, tRNA availability, and translation factor abundance, as well as the relationship between the overall rates of initiation, elongation, and termination. Stress, antibiotics, and genetic perturbations affecting composition and properties of translation machinery can alter the ribosome positional distribution dramatically. Here, we offer a computational protocol for analyzing positional distribution profiles using ribosome profiling (Ribo-Seq) data. The protocol uses papolarity, a new Python toolkit for the analysis of transcript-level short read coverage profiles. For a single sample, for each transcript papolarity allows for computing the classic polarity metric which, in the case of Ribo-Seq, reflects ribosome positional preferences. For comparison versus a control sample, papolarity estimates an improved metric, the relative linear regression slope of coverage along transcript length. This involves de-noising by profile segmentation with a Poisson model and aggregation of Ribo-Seq coverage within segments, thus achieving reliable estimates of the regression slope. The papolarity software and the associated protocol can be conveniently used for Ribo-Seq data analysis in the command-line Linux environment. Papolarity package is available through Python pip package manager. The source code is available at https://github.com/autosome-ru/papolarity .
The authors would like to correct Figure 3, panel J, in which the rightmost upper image of SA-β-gal stained 293T cells following short hairpin RNA (shRNA)-mediated knockdown of RRAS2 with sh769 (RRAS2-KD-sh769) was inadvertently, and due to a labeling error, taken from the same original source image presented in the middle upper panel, which shows increased SA-β-gal activity following RRAS2 knockdown by a different shRNA (sh646).This correction does not affect any of the conclusions of the article.The corrected image representative of RRAS2-KD-sh769 is provided below, and Figure 3 has been updated in the article online.
Long noncoding RNAs (lncRNAs) constitute the majority of transcripts in the mammalian genomes, and yet, their functions remain largely unknown. As part of the FANTOM6 project, we systematically knocked down the expression of 285 lncRNAs in human dermal fibroblasts and quantified cellular growth, morphological changes, and transcriptomic responses using Capped Analysis of Gene Expression (CAGE). Antisense oligonucleotides targeting the same lncRNAs exhibited global concordance, and the molecular phenotype, measured by CAGE, recapitulated the observed cellular phenotypes while providing additional insights on the affected genes and pathways. Here, we disseminate the largest-to-date lncRNA knockdown data set with molecular phenotyping (over 1000 CAGE deep-sequencing libraries) for further exploration and highlight functional roles for ZNF213-AS1 and lnc-KHDC3L-2 .
Knowledge of mechanisms responsible for mutagenesis of adult stem cells is crucial to track genomic alterations that may affect cell renovation and provoke malignant cell transformation. Mutations in regulatory regions are widely studied nowadays, though mostly in cancer. In this study, we decomposed the mutation signature of adult stem cells, mapped the corresponding mutations into transcription factor binding regions, and assessed mutation frequency in sequence motif occurrences. We found binding sites of C/EBP transcription factors strongly enriched with [C>T]G mutations within the core CG dinucleotide related to deamination of the methylated cytosine. This effect was also exhibited in related cancer samples. Structural modeling predicted enhanced CEBPB binding to the consensus sequence with the [C>T]G mismatch, which was then confirmed in the direct experiment. We propose that it is the enhanced binding of C/EBPs that shields C>T transitions from DNA repair and leads to selective accumulation of the [C>T]G mutations within binding sites.