Spatial barcoding-based transcriptomic (ST) data require deconvolution for cellular-level downstream analysis. Here we present SDePER, a hybrid machine learning and regression method to deconvolve ST data using reference single-cell RNA sequencing (scRNA-seq) data. SDePER tackles platform effects between ST and scRNA-seq data, ensuring a linear relationship between them while addressing sparsity and spatial correlations in cell types across capture spots. SDePER estimates cell-type proportions, enabling enhanced resolution tissue mapping by imputing cell-type compositions and gene expressions at unmeasured locations. Applications to simulated data and four real datasets showed SDePER’s superior accuracy and robustness over existing methods.
Histone modifications (HMs) are pivotal in various biological processes, including transcription, replication, and DNA repair, significantly impacting chromatin structure. These modifications underpin the molecular mechanisms of cell-type-specific gene expression and complex diseases. However, annotating HMs across different cell types solely using experimental approaches is impractical due to cost and time constraints. Herein, we present dHICA (deep histone imputation using chromatin accessibility), a novel deep learning framework that integrates DNA sequences and chromatin accessibility data to predict multiple HM tracks. Employing the transformer architecture alongside dilated convolutions, dHICA boasts an extensive receptive field and captures more cell-type-specific information. dHICA outperforms state-of-the-art baselines and achieves superior performance in cell-type-specific loci and gene elements, aligning with biological expectations. Furthermore, dHICA's imputations hold significant potential for downstream applications, including chromatin state segmentation and elucidating the functional implications of SNPs (Single Nucleotide Polymorphisms). In conclusion, dHICA serves as a valuable tool for advancing the understanding of chromatin dynamics, offering enhanced predictive capabilities and interpretability.
Histone modifications (HMs) play a pivot role in various biological processes, including transcription, replication and DNA repair, significantly impacting chromatin structure. These modifications underpin the molecular mechanisms of cell-specific gene expression and complex diseases. However, annotating HMs across different cell types solely using experimental approaches is impractical due to cost and time constraints. Herein, we present dHICA (discriminative histone imputation using chromatin accessibility), a novel deep learning framework that integrates DNA sequences and chromatin accessibility data to predict multiple HM tracks. Employing the Transformer architecture alongside dilated convolutions, dHICA boasts an extensive receptive field and captures more cell-type-specific information. dHICA not only outperforms state-of-the-art baselines but also achieves superior performance in cell-specific loci and gene elements, aligning with biological expectations. Furthermore, dHICA’s imputations hold significant potential for downstream applications, including chromatin state segmentation and elucidating the functional implications of SNPs. In conclusion, dHICA serves as an invaluable tool for advancing the understanding of chromatin dynamics, offering enhanced predictive capabilities and interpretability.### Competing Interest StatementThe authors have declared no competing interest.
Run-on and sequencing assays like GRO-seq, PRO-seq, and ChRO-seq allow for joint profiling of transcription activity of transcriptional regulatory elements (TREs), i.e., promoters and active enhancers, and target genes. Variation in biological conditions, such as treated vs. control, results in changes in the activity of transcription factors (TFs), which induces concerted changes in TREs and target genes. By modeling the differences between two biological conditions, we developed the computational pipeline known as tfTarget that predicts a set of putative TREs and target genes responding to each TF under the biological condition of interest. In this chapter, we demonstrate the use of the new web-based tfTarget in mapping transcription regulation using run-on sequencing data.
Genome-wide association studies (GWAS), particularly designed with thousands and thousands of single-nucleotide polymorphisms (SNPs) (big p) genotyped on tens of thousands of subjects (small n), are encountered by a major challenge of p << n. Although the integration of longitudinal information can significantly enhance a GWAS's power to comprehend the genetic architecture of complex traits and diseases, an additional challenge is generated by an autocorrelative process. We have developed several statistical models for addressing these two challenges by implementing dimension reduction methods and longitudinal data analysis. To make these models computationally accessible to applied geneticists, we wrote an R package of computer software, HiGwas, designed to analyze longitudinal GWAS datasets. Functions in the package encompass single SNP analyses, significance-level adjustment, preconditioning and model selection for a high-dimensional set of SNPs. HiGwas provides the estimates of genetic parameters and the confidence intervals of these estimates. We demonstrate the features of HiGwas through real data analysis and vignette document in the package.
Quantitative trait locus (QTL) mapping has been used as a powerful tool for inferring the complexity of the genetic architecture that underlies phenotypic traits. This approach has shown its unique power to map the developmental genetic architecture of complex traits by implementing longitudinal data analysis. Here, we introduce the R package Funmap2 based on the functional mapping framework, which integrates prior biological knowledge into the statistical model. Specifically, the functional mapping framework is engineered to include longitudinal curves that describe the genetic effects and the covariance matrix of the trait of interest. Funmap2 chooses the type of longitudinal curve and covariance matrix automatically using information criteria. Funmap2 is available for download at https://github.com/wzhy2000/Funmap2.
The past two decades have witnessed a revolution in identifying genetic risk factors underlying diseases and complex traits using genome-wide association studies(GWAS)(Risch and Merikangas,1996;Hirschhorn and Daly,2005;Altshuler et al.,2008).Together with advanced high-throughput technologies for genotyping and
Functional Mapping is a popular statistical method in QTL mapping studies for longitudinal data. The threshold for declaring statistical significance of a QTL is commonly obtained through permutation tests, which can be time consuming. To improve the computational efficiency of a permutation test of mixture models used in Functional Mapping, we first quantified the correlation between QTL and longitudinal data, using a curve clustering method. Then, the QTLs which are highly correlated with the outcome were computed in the improved permutation tests. As a result, it reduces the amount of computation in permutation tests and speeds up the computation for Functional Mapping analysis. Simulation studies and real data analysis were conducted to demonstrate that the proposed approach can greatly improve the computational efficiency of QTL mapping without loss of accuracy.