The sequence of the human genome provides a foundation for understanding cellular processes in health and disease1. The organisation of this primary genetic information into cell-specific structure and function is critical to understanding the cell type-specific interpretation and execution of the genome. Epigenetic processes are essential for packaging and higher-level functional organisation of the genome, and changes therein are increasingly recognised as contributors to human disease. Building on primary data generated by multinational consortia, the International Human Epigenome Consortium2 (IHEC) has uniformly processed a collection of more than 2000 comprehensive human reference epigenomes, collectively referred to as EpiATLAS. This effort involved the development of standardised molecular and bioinformatics protocols, metadata models, and analytical tools to manage, integrate, display, and share vast amounts of epigenomic data. This includes the creation of a publicly available Epigenome Reference Registry, which provides a system for accessing protected human subject datasets and facilitates open searching of de-identified samples and experimental data. The integrated EpiATLAS ecosystem and its comprehensive human reference epigenome maps provide an unprecedented resource for the biosciences, expanding the annotated epigenomic landscape while uncovering previously unappreciated relationships among regulatory layers and revealing how epigenetic inputs underpin fundamental cellular functions and disease associations.
The International Human Epigenome Consortium has generated thousands of datasets profiling transcription factor binding, histone modification, and DNA accessibility. Existing segmentation and annotation approaches typically produce cell type-specific regulatory maps, which have become increasingly difficult to maintain and apply as the number of profiled cell types has grown. Here, we apply epigenome-ssm, a continuous state-space modeling framework, to generate a unified, interpretable representation of chromatin states across thousands of human epigenomes. Using 9,539 histone modification signal tracks from 1,698 epigenomes, epigenome-ssm produces 33 continuous chromatin state features that compactly capture regulatory programs across cell types. These features distinguish canonical activities such as promoters, enhancers, transcription, and heterochromatin, while also encoding cell type-specific regulatory patterns. Compared with alternative pan-cell type annotation methods, the continuous features achieve superior or comparable predictive performance for gene expression, enhancer activity, and evolutionary conservation, despite using fewer dimensions. The model effectively captures both broad and lineage-specific regulatory programs, linking chromatin states to gene expression and functional annotations. Additionally, a derived conservation-associated activity score (SSM-CAAS) highlights genomic regions enriched for disease-associated variants, demonstrating utility for interpreting noncoding variation. Continuous pan-cell type chromatin state features provide a compact, expressive, and biologically informative representation of the human epigenome. This framework improves integration and interpretation of large-scale epigenomic data, enables accurate prediction of genomic function, and facilitates identification of regulatory elements relevant to disease. The resulting resource offers a scalable foundation for downstream analyses of gene regulation and genetic variation.
Large-scale epigenomic datasets such as histone modifications and DNA accessibility have greatly advanced our understanding of genomic function. However, these measurements often suffer from noise, batch effects and irreproducibility. Epigenome imputation has emerged as a promising solution to these challenges. These methods integrate patterns across experiments, cell types, and genomic loci to predict the results of experiments, yielding predictions that often surpass observed data in quality. Thus, researchers increasingly leverage imputation for denoising data prior to downstream analysis. However, existing methods for imputation-based denoising have significant limitations. Here, we propose CANDI (Confidence-Aware Neural Denoising Imputer), a method for epigenome imputation that (1) predicts raw counts and handles experiment-specific covariates such as sequencing depth, (2) can (optionally) incorporate information from a low-quality existing experiment when predicting a target without retraining, and (3) outputs a calibrated measure of uncertainty. This approach is enabled using a Transformer model with self-supervised learning (SSL) training. ### Competing Interest Statement The authors have declared no competing interest.
The fourth and final phase of the ENCODE consortium has newly profiled epigenetic activity in hundreds of human tissues. Chromatin state annotations created by segmentation and genome annotation (SAGA) methods such as Segway have emerged as the predominant integrative summary of such data sets. Here, we present the ENCODE4 Catalog of Segway Annotations, a set of sample-specific genome-wide chromatin state annotations of 234 human biosamples inferred from 1,794 genomics experiments. This catalog identifies genomic elements, accurately captures cell type-specific regulatory patterns, and facilitates discovery of elements involved in phenotype and disease.
Single-cell Hi-C (scHi-C) enables the study of 3D genome organization at the resolution of individual cells and cell types that cannot be isolated for bulk profiling. However, the extreme sparsity of scHi-C data presents major challenges, particularly in recovering cell-type-specific 3D structures when only a small number of cells are available. We introduce contactVI, a method that combines the strengths of graph-based models and variational autoencoders (VAEs) to account for spatial dependencies in noisy chromatin interaction data and effectively denoise them. On simulated data, contactVI outperforms existing imputation methods in recovering Hi-C contact maps at both the single-cell and cell-type levels. On real datasets, contactVI performs comparably to or better than other graph-based methods across different resolutions. When applied to jointly profiled single-cell Hi-C and RNA-seq data, contactVI successfully recovers the expected association between genome compartmentalization and gene expression. contactVI is available on GitHub: https://github.com/nedashokraneh/contactVI
Understanding the mechanistic basis of genetic disease requires annotating the regulatory elements in the human genome. To this end, the International Human Epigenome Consortium (IHEC) has generated thousands of epigenomic datasets--including ChIP-seq, DNase-seq, and ATAC-seq--that measure various biochemical activities in the genome, including transcription factor binding, histone modification, and DNA accessibility. Currently, the predominant methods for integrating these data sets to annotate regulatory elements are segmentation and genome annotation (SAGA) algorithms such as ChromHMM and Segway. The majority of annotations generated by these methods are cell type-specific. However, as the number of profiled cell types has grown into the thousands, using thousands of cell type-specific chromatin state annotations proves undesirable for many applications. Therefore, recently, researchers have sought a single unified annotation of regulatory elements across all cell types, known as a "pan-cell type" annotation. Here, we present a pan-cell type annotation that summarizes all IHEC epigenomes using the recently-developed method, epigenome-ssm. This pan-cell type annotation comprises 33 genome-wide chromatin state feature signal tracks, each of which captures a regulatory program driving genomic activity in one or more cell types. We show that these feature maps constitute an intuitive and visualizable summary of epigenomic data. ### Competing Interest Statement The authors have declared no competing interest.
MOTIVATION:The genome-wide chromosome conformation capture assay Hi-C is widely used to study chromatin 3D structures and their functional implications. Read counts from Hi-C indicate the strength of chromatin contact between each pair of genomic loci. These read counts are heteroskedastic: that is, a difference between the interaction frequency of 0 and 100 is much more significant than a difference between the interaction frequency of 1000 and 1100. This property impedes visualization and downstream analysis because it violates the Gaussian variable assumption of many computational tools. Thus heuristic transformations aimed at stabilizing the variance of signals like the shifted-log transformation are typically applied to data before its visualization and inputting to models with Gaussian assumption. However, such heuristic transformations cannot fully stabilize the variance because of their restrictive assumptions about the mean-variance relationship in the data. RESULTS:Here, we present VSS-Hi-C, a data-driven variance stabilization method for Hi-C data. We show that VSS-Hi-C signals have a unit variance improving visualization of Hi-C, for example in heatmap contact maps. VSS-Hi-C signals also improve the performance of subcompartment callers relying on Gaussian observations. VSS-Hi-C is implemented as an R package and can be used for variance stabilization of different genomic and epigenomic data types with two replicates available. AVAILABILITY AND IMPLEMENTATION:https://github.com/nedashokraneh/vssHiC.
With the goal of mapping genomic activity, international projects have recently measured epigenetic activity in hundreds of cell and tissue types. Chromatin state annotations produced by segmentation and genome annotation (SAGA) methods have emerged as the predominant way to summarize these epigenomic data sets in order to annotate the genome. These chromatin state annotations are essential for many genomic tasks, including identifying active regulatory elements and interpreting disease-associated genetic variation. However, despite the widespread applications of SAGA methods, no principled approach exists to evaluate the statistical significance of chromatin state assignments. Here, we propose the first method for assigning calibrated confidence scores to chromatin state annotations. Toward this goal, we performed a comprehensive evaluation of the reproducibility of the two most widely used existing SAGA methods, ChromHMM and Segway. We found that their predictions are frequently irreproducible. For example, when applying the same SAGA method on two sets of experimental replicates, 27%-69% of predicted enhancers fail to replicate. This suggests that a substantial fraction of predicted elements in existing chromatin state annotations cannot be relied upon. To remedy this problem, we introduce SAGAconf, a method for assigning a measure of confidence (r-value) to chromatin state annotations. SAGAconf works with any SAGA method and assigns an r-value to each genomic bin of a chromatin state annotation that represents the probability that the label of this bin will be reproduced in a replicated experiment. Thus, SAGAconf allows a researcher to select only the reliable predictions from a chromatin annotation for use in downstream analyses.
Background Segmentation and genome annotations (SAGA) methods such as ChromHMM and Segway are widely to annotate chromatin states in the genome. These algorithms take as input a collection of genomics datasets, partition the genome, and assign a label to each segment such that positions with the same label have similar patterns in the input data. SAGA methods output an human-interpretable summary of the genome by labeling every genomic position with its annotated activity such as Enhancer, Transcribed, etc. Chromatin state annotations are essential for many genomic tasks, including identifying active regulatory elements and interpreting disease-associated genetic variation. However, despite the widespread applications of SAGA methods, no principled approach exists to evaluate the statistical significance of SAGA state assignments. Results Towards the goal of producing robust chromatin state annotations, we performed a comprehensive evaluation of the reproducibility of SAGA methods. We show that SAGA annotations exhibit a large degree of disagreement, even when run with the same method on replicated data sets. This finding suggests that there is significant risk to using SAGA chromatin state annotations. To remedy this problem, we introduce SAGAconf, a method for assigning a measure of confidence (r-value) to SAGA annotations. This r-value is assigned to each genomic bin of a SAGA annotation and represents the probability that the label of this bin will be reproduced in a replicated experiment. This process is analogous to irreproducible discovery rate (IDR) analysis that is commonly used for ChIP-seq peak calling and related tasks. Thus SAGAconf allows a researcher to select only the reliable parts of a SAGA annotation for use in downstream analyses. SAGAconf r-values provide accurate confidence estimates of SAGA annotations, allowing researchers to filter out unreliable elements and remove doubt in those that stand up to this scrutiny.
Towards the goal of identifying functional elements in the human genome, the fourth and final phase of the ENCODE consortium has newly profiled hundreds of human tissues using sequencing-based measurements of genomic activity such as ChIP-seq measures of transcription factor binding and histone modification. Chromatin state annotations created by segmentation and genome annotation (SAGA) methods such as Segway have emerged as the predominant integrative summary of such epigenomic data sets. Here, we present the ENCODE4 catalog of Segway annotations, a set of sample-specific genome-wide Segway chromatin state annotations for 234 ENCODE human biosamples inferred from 1,794 functional genomics experiments. We define an updated vocabulary of chromatin state terms that includes patterns of activity present only in a subset of samples or identified only with rarely-performed assays. We show that these ENCODE4 Segway annotations accurately capture both general and cell-type-specific regulatory patterns, and do so with substantially improved sensitivity relative to prior large-scale chromatin annotation sets. This catalog facilitates the downstream discovery of regulatory mechanisms which underlie diseases and traits identified by genome-wide association studies.
Recent advances in single-cell Hi-C (scHi-C) assays allow studying the chromatin conformation at the resolution of a single cell or a cluster of cells. A key question is to identify changes in the contact strength between two cell types, known as differential chromatin contacts (DCCs). While existing statistical methods can identify changes in contact strength in bulk Hi-C data, these methods cannot be effectively applied to scHi-C data due to its severe sparsity. Thus it is necessary to develop methods for identifying differential chromatin contacts in scHi-C data. Recently-developed scHi-C imputation approaches can mitigate the issue of sparsity. We propose an approach for identifying differential chromatin contacts using these imputation approaches. We build upon the existing SnapHiC-D method by replacing its imputation step with recent learning-based imputation approaches. We show that, via analysis of real scHi-C datasets with different coverages and at different resolutions, imputation approaches that consider the spatial correlation between bin pairs, Higashi, and random walk with restart, outperform other approaches. Furthermore, we show that careful considerations are needed when imputation is done in preprocessing steps as it may invalidate downstream statistical approaches. Finally, our results indicate that model-based imputations greatly improve performance when analyzing chromatin contacts at moderate resolution (100kb); however, current imputation approaches are inefficient in terms of both accuracy and computational complexity when being applied to high-resolution scHi-C resolution (10kb).
Drug repurposing can accelerate the identification of effective compounds for clinical use against SARS-CoV-2, with the advantage of pre-existing clinical safety data and an established supply chain. RNA viruses such as SARS-CoV-2 manipulate cellular pathways and induce reorganization of subcellular structures to support their life cycle. These morphological changes can be quantified using bioimaging techniques. In this work, we developed DEEMD: a computational pipeline using deep neural network models within a multiple instance learning framework, to identify putative treatments effective against SARS-CoV-2 based on morphological analysis of the publicly available RxRx19a dataset. This dataset consists of fluorescence microscopy images of SARS-CoV-2 non-infected cells and infected cells, with and without drug treatment. DEEMD first extracts discriminative morphological features to generate cell morphological profiles from the non-infected and infected cells. These morphological profiles are then used in a statistical model to estimate the applied treatment efficacy on infected cells based on similarities to non-infected cells. DEEMD is capable of localizing infected cells via weak supervision without any expensive pixel-level annotations. DEEMD identifies known SARS-CoV-2 inhibitors, such as Remdesivir and Aloxistatin, supporting the validity of our approach. DEEMD can be explored for use on other emerging viruses and datasets to rapidly identify candidate antiviral treatments in the future}. Our implementation is available online at https://www.github.com/Sadegh-Saberian/DEEMD
Motivation With the advancement of sequencing technologies, genomic data sets are constantly being expanded by high volumes of different data types. One recently introduced data type in genomic science is genomic signals, which are usually short-read coverage measurements over the genome. An example of genomic signals is Epigenomic marks which are utilized to locate functional and nonfunctional elements in genome annotation studies. To understand and evaluate the results of such studies, one needs to understand and analyze the characteristics of the input data. Results SigTools is an R-based genomic signals visualization package developed with two objectives: 1) to facilitate genomic signals exploration in order to uncover insights for later model training, refinement, and development by including distribution and autocorrelation plots. 2) to enable genomic signals interpretation by including correlation , and aggregation plots. Moreover, Sigtools also provides text-based descriptive statistics of the given signals which can be practical when developing and evaluating learning models. We also include results from 2 case studies. The first examines several previously studied genomic signals called histone modifications. This use case demonstrates how SigTools can be beneficial for satisfying scientists’ curiosity in exploring and establishing recognized datasets. The second use case examines a dataset of novel chromatin state features which are novel genomic signals generated by a learning model. This use case demonstrates how SigTools can assist in exploring the characteristics and behavior of novel signals towards their interpretation. In addition, our corresponding web application, SigTools-Shiny, extends the accessibility scope of these modules to people who are more comfortable working with graphical user interfaces instead of command-line tools. Availability SigTools source code, installation guide, and manual is available on http://github.com/shohre73 . Contact shohre_masoumi@sfu.ca
Semi-automated genome annotation (SAGA) methods are widely used to understand genome activity and gene regulation. These methods take as input a set of sequencing-based assays of epigenomic activity (such as ChIP-seq measurements of histone modification and transcription factor binding), and output an annotation of the genome that assigns a chromatin state label to each genomic position. Existing SAGA methods have several limitations caused by the discrete annotation framework: such annotations cannot easily represent varying strengths of genomic elements, and they cannot easily represent combinatorial elements that simultaneously exhibit multiple types of activity. To remedy these limitations, we propose an annotation strategy that instead outputs a vector of chromatin state features at each position rather than a single discrete label. Continuous modeling is common in other fields, such as in topic modeling of text documents. We propose a method, epigenome-ssm, that uses a Kalman filter state space model to efficiently annotate the genome with chromatin state features. We show that chromatin state features from epigenome-ssm are more useful for several downstream applications than both continuous and discrete alternatives, including their ability to identify expressed genes and enhancers. Therefore, we expect that these continuous chromatin state features will be valuable reference annotations to be used in visualization and downstream analysis.
The availability of thousands of assays of epigenetic activity necessitates compressed representations of these data sets that summarize the epigenetic landscape of the genome. Until recently, most such representations were celltype specific, applying to a single tissue or cell state. Recently, neural networks have made it possible to summarize data across tissues to produce a pan-celltype representation. In this work, we propose Epi-LSTM, a deep long short-term memory (LSTM) recurrent neural network autoencoder to capture the long-term dependencies in the epigenomic data. The latent representations from Epi-LSTM capture a variety of genomic phenomena, including gene-expression, promoter-enhancer interactions, replication timing, frequently interacting regions and evolutionary conservation. These representations outperform existing methods in a majority of cell-types, while yielding smoother representations along the genomic axis due to their sequential nature.
Despite the availability of chromatin conformation capture experiments, discerning the relationship between the 1D genome and 3D conformation remains a challenge, which limits our understanding of their affect on gene expression and disease. We propose Hi-C-LSTM, a method that produces low-dimensional latent representations that summarize intra-chromosomal Hi-C contacts via a recurrent long short-term memory neural network model. We find that these representations contain all the information needed to recreate the observed Hi-C matrix with high accuracy, outperforming existing methods. These representations enable the identification of a variety of conformation-defining genomic elements, including nuclear compartments and conformation-related transcription factors. They furthermore enable in-silico perturbation experiments that measure the influence of cis-regulatory elements on conformation. Despite the availability of chromatin conformation capture experiments, discerning the relationship between the 1D genome and 3D conformation remains a challenge. Here, the authors propose a method that produces low-dimensional latent representations that summarize intra-chromosomal Hi-C contacts.
Artificial intelligence (AI) models based on deep learning now represent the state of the art for making functional predictions in genomics research. However, the underlying basis on which predictive models make such predictions is often unknown. For genomics researchers, this missing explanatory information would frequently be of greater value than the predictions themselves, as it can enable new insights into genetic processes. We review progress in the emerging area of explainable AI (xAI), a field with the potential to empower life science researchers to gain mechanistic insights into complex deep learning models. We discuss and categorize approaches for model interpretation, including an intuitive understanding of how each approach works and their underlying assumptions and limitations in the context of typical high-throughput biological datasets.
Segmentation and genome annotation (SAGA) algorithms are widely used to understand genome activity and gene regulation. These algorithms take as input epigenomic datasets, such as chromatin immunoprecipitation-sequencing (ChIP-seq) measurements of histone modifications or transcription factor binding. They partition the genome and assign a label to each segment such that positions with the same label exhibit similar patterns of input data. SAGA algorithms discover categories of activity such as promoters, enhancers, or parts of genes without prior knowledge of known genomic elements. In this sense, they generally act in an unsupervised fashion like clustering algorithms, but with the additional simultaneous function of segmenting the genome. Here, we review the common methodological framework that underlies these methods, review variants of and improvements upon this basic framework, catalogue existing large-scale reference annotations, and discuss the outlook for future work.
Segmentation and genome annotation (SAGA) algorithms are widely used to understand genome activity and gene regulation. These algorithms take as input epigenomic datasets, such as chromatin immunoprecipitation-sequencing (ChIP-seq) measurements of histone modifications or transcription factor binding. They partition the genome and assign a label to each segment such that positions with the same label exhibit similar patterns of input data. SAGA algorithms discover categories of activity such as promoters, enhancers, or parts of genes without prior knowledge of known genomic elements. In this sense, they generally act in an unsupervised fashion like clustering algorithms, but with the additional simultaneous function of segmenting the genome. Here, we review the common methodological framework that underlies these methods, review variants of and improvements upon this basic framework, and discuss the outlook for future work. This review is intended for those interested in applying SAGA methods and for computational researchers interested in improving upon them.