In hepatocellular carcinoma (HepG2), aberrant histone modifications are closely linked to long non-coding RNA (lncRNA) expression. However, existing computational models lack physical interpretability at specific promoter coordinates. To address this, we developed a position-specific statistical scoring model based on adjacent and next-adjacent nucleotide frequencies. We trained two independent, position-specific matrices representing increased and decreased modification states across 600 bp promoter windows centered on the true signal summits. Finally, ten-fold cross-validation revealed that significant energy differences between sequences with increased and decreased histone signals enable excellent classification performance. These results indicted a strong correlation between the total energy of local DNA structures and histone modification signal.
In hepatocellular carcinoma (HCC), aberrant histone modifications are linked to the dysregulation of long non-coding RNA (lncRNA) expression. Although existing computational models can accurately predict some associations, they lack deep physical interpretability. We constructed an energy model based on the physical principle that energy determines molecular structure. Total DNA segment energy was calculated by summing adjacent trinucleotide interaction energies and applied to analyze 11 key histone modifications in HCC, specifically within lncRNA promoter regions where modification signals were increased or decreased. Finally, ten-fold cross-validation revealed that significant energy differences between sequences with increased and decreased histone signals enable excellent classification performance. These results indicted a strong correlation between the total energy of local DNA structures and histone modification signal. Furthermore, introducing longer k-mers led to computational redundancy without a consistent improvement, confirming that the trinucleotide model most effectively acquires the local DNA structural changes associated with histone modification levels. Our model can effectively distinguish DNA sequences associated with different histone modification levels from a physical energy perspective. This model serves as an interpretable tool for epigenetic research while providing a new understanding a new perspective for understanding the dysregulation of lncRNA expression in HCC.
Stomach adenocarcinoma (STAD) has high incidence and mortality rates. Long non-coding RNAs (lncRNAs) and angiogenesis are closely related to the pathogenesis and metastasis of STAD. Recently, emerging evidence demonstrated that DNA methylation plays crucial roles in the development of STAD. This study explored the relationship between DNA methylation and the abnormal expression of angiogenesis-related lncRNAs (ARlncRNAs) in stomach adenocarcinoma, aiming to identify prognostic biomarkers. Moreover, a Cox analysis and Lasso regression were used to establish an ARlncRNA feature set related to angiogenesis. The prognostic model was evaluated by using a Kaplan-Meier (KM) analysis, ROC curves, and nomograms. Based on the identified 18 key ARlncRNAs, a prognostic predictive model was constructed. In addition, a specific ARlncRNA with abnormal methylation in the model, LINC00511, showed significant differences in expression and methylation across different subgroups. The methylation and expression of LINC00511 were analyzed by a correlation and co-expression analysis. The correlation analysis indicated that promoter methylation may improve LINC00511 expression. Further analysis found 355 mRNAs co-expressed with LINC00511 which may interact with 6 miRNAs to regulate target gene expression. The abnormal methylation of LINC00511 could significantly contribute to the progression of stomach adenocarcinoma.
Overall cancer hypomethylation had been identified in the past, but it is not clear exactly which hypomethylation site is the more important for the occurrence of cancer. To identify key hypomethylation sites, we studied the effect of hypomethylation in twelve regions on gene expression in colon adenocarcinoma (COAD). The key DNA methylation sites of cg18949415, cg22193385 and important genes of C6orf223, KRT7 were found by constructing a prognostic model, survival analysis and random combination prediction a series of in-depth systematic calculations and analyses, and the results were validated by GEO database, immune microenvironment, drug and functional enrichment analysis. Based on the expression values of C6orf223, KRT7 genes and the DNA methylation values of cg18949415, cg22193385 sites, the least diversity increment algorithm were used to predict COAD and normal sample. The 100 % reliability and 97.12 % correctness of predicting tumor samples were obtained in jackknife test. Moreover, we found that C6orf223 gene, cg18949415 site play a more important role than KRT7 gene, cg22193385 site in COAD. In addition, we investigate the impact of key methylation sites on three-dimensional chromatin structure. Our results will be help for experimental studies and may be an epigenetic biomarker for COAD.
Aberrant DNA methylation plays a crucial role in breast cancer progression by regulating gene expression. However, the regulatory pattern of DNA methylation in long noncoding RNAs (lncRNAs) for breast cancer remains unclear. In this study, we integrated gene expression, DNA methylation, and clinical data from breast cancer patients included in The Cancer Genome Atlas (TCGA) database. We examined DNA methylation distribution across various lncRNA categories, revealing distinct methylation characteristics. Through genome-wide correlation analysis, we identified the CpG sites located in lncRNAs and the distally associated CpG sites of lncRNAs. Functional genome enrichment analysis, conducted through the integration of ENCODE ChIP-seq data, revealed that differentially methylated CpG sites (DMCs) in lncRNAs were mostly located in promoter regions, while distally associated DMCs primarily acted on enhancer regions. By integrating Hi-C data, we found that DMCs in enhancer and promoter regions were closely associated with the changes in three-dimensional chromatin structures by affecting the formation of enhancer-promoter loops. Furthermore, through Cox regression analysis and three machine learning models, we identified 11 key methylation-driven lncRNAs (DIO3OS, ELOVL2-AS1, MIAT, LINC00536, C9orf163, AC105398.1, LINC02178, MILIP, HID1-AS1, KCNH1-IT1, and TMEM220-AS1) that were associated with the survival of breast cancer patients and constructed a prognostic risk scoring model, which demonstrated strong prognostic performance. These findings enhance our understanding of DNA methylation's role in lncRNA regulation in breast cancer and provide potential biomarkers for diagnosis.
The discovery of key epigenetic modifications in cancer is of great significance for the study of disease biomarkers. Through the mining of epigenetic modification data relevant to cancer, some researches on epigenetic modifications are accumulating. In order to make it easier to integrate the effects of key epigenetic modifications on the related cancers, we established CancerMHL (http://www.positionprediction.cn/), which provide key DNA methylation, histone modifications and lncRNAs as well as the effect of these key epigenetic modifications on gene expression in several cancers. To facilitate data retrieval, CancerMHL offers flexible query options and filters, allowing users to access specific key epigenetic modifications according to their own needs. In addition, based on the epigenetic modification data, three online prediction tools had been offered in CancerMHL for users. CancerMHL will be a useful resource platform for further exploring novel and potential biomarkers and therapeutic targets in cancer.Database URL: http://www.positionprediction.cn/
IntroductionLong non-coding RNAs (lncRNAs) play crucial roles in genetic markers, genome rearrangement, chromatin modifications, and other biological processes. Increasing evidence suggests that lncRNA functions are closely related to their subcellular localization. However, the distribution of lncRNAs in different subcellular localizations is imbalanced. The number of lncRNAs located in the nucleus is more than ten times that in the exosome.MethodsIn this study, we propose a new oversampling method to construct a predictive dataset and develop a predictive model called LncSTPred. This model improves the Adaboost algorithm for subcellular localization prediction using 3-mer, 3-RF sequence, and minimum free energy structure features.Results and DiscussionBy using our improved Adaboost algorithm, better prediction accuracy for lncRNA subcellular localization was obtained. In addition, we evaluated feature importance by using the F-score and analyzed the influence of highly relevant features on lncRNAs. Our study shows that the ANA features may be a key factor for predicting lncRNA subcellular localization, which correlates with the composition of stems and loops in the secondary structure of lncRNAs.
As the direct recipients of plant abscisic acid (ABA), pyrabactin resistance/pyrabactin resistance-like/regulatory component of ABA receptor proteins (PYR/PYL/RCAR; hereinafter called PYLs) play pivotal roles in plant coercive responses. However, the PYL genes in Acer palmatum have yet to be investigated. In this study, using a genome search method, we identified 14 A. palmatum PYL genes (ApPYLs) and clustered them into three clades based on the phylogenetic, gene structure, and conserved motif analyses. The ApPYLs were dispersed on eight chromosomes of A. palmatum; cis-acting element analysis indicated that the ApPYLs were involved in biological processes, including resistance, hormone regulation, and growth. Gene expression profiling of cold-treated A. palmatum plants revealed that ApPYL1 and ApPYL2 may be related to cold stress response. ApPYL1 and ApPYL2 interacted directly with ApMYB44, and the ApPYL1, ApPYL2, and ApMYB44 genes were found to enhance freezing tolerance in Arabidopsis. These findings illustrate that ApPYL1 and ApPYL2 respond to the cold signalling pathway by interacting with ApMYB44 and provide a reference for improving cold resistance of plants.
Acute myeloid leukemia (AML) is a rare tumor that invades the blood and bone marrow, it is rapidly progressive, highly aggressive, and difficult to cure. Studies have shown that long non-coding RNA (lncRNA) and ferroptosis play important roles in AML. However, few studies have been done on ferroptosis-related lncRNA for AML. To investigate the role of ferroptosis-related lncRNA in AML prognosis, we screened the differentially expressed genes related to ferroptosis and lncRNA. Ferroptosis-related lncRNA associated with AML prognosis was obtained by Pearson correlation analysis. By using univariate Cox analysis, least absolute shrinkage and selection operator (LASSO) analysis, and multivariate Cox analysis, the ten prognostic genes were used for constructing the prognostic model. The model was then validated using a Kaplan-Meier analysis and Cox regression analysis. The ROC results have shown that the model could better predict AML survival. We identified some mutated genes that may affect the poor prognosis based on the somatic mutation analysis. The enrichment pathway analysis of prognostic genes revealed that these genes were mainly enriched in some immune pathways and cancer pathways. By immune infiltration analysis, we found that high-risk patients may respond better to immunotherapy.
Aim: The present study was designed to investigate the coregulatory effects of multiple histone modifications (HMs) on gene expression in lung adenocarcinoma (LUAD). Materials & methods: Ten histones for LUAD were analyzed using ChIP-seq and RNA-seq data. An innovative computational method is proposed to quantify the coregulatory effects of multiple HMs on gene expression to identify strong coregulatory genes and regions. This method was applied to explore the coregulatory mechanisms of key ferroptosis-related genes in LUAD. Results: Nine strong coregulatory regions were identified for six ferroptosis-related genes with diverse coregulatory patterns (CA9, PGD, CDKN2A, PML, OTUB1 and NFE2L2). Conclusion: This quantitative method could be used to identify important HM coregulatory genes and regions that may be epigenetic regulatory targets in cancers. A new computational method is proposed to quantify the effects of multiple histone modifications in coregulating gene expression. This work was designed to study the coregulation of ferroptosis-related genes in lung adenocarcinoma and determine key coregulatory genes and regions.
Acute myeloid leukemia (AML) is an aggressive malignancy characterized by challenges in treatment, including drug resistance and frequent relapse. Recent research highlights the crucial roles of tumor microenvironment (TME) in assisting tumor cell immune escape and promoting tumor aggressiveness. This study delves into the interplay between AML and TME. Through the exploration of potential driver genes, we constructed an AML prognostic index (AMLPI). Cross-platform data and multi-dimensional internal and external validations confirmed that the AMLPI outperforms existing models in terms of areas under the receiver operating characteristic curves, concordance index values, and net benefits. High AMLPIs in AML patients were indicative of unfavorable prognostic outcomes. Immune analyses revealed that the high-AMLPI samples exhibit higher expression of HLA-family genes and immune checkpoint genes (including PD1 and CTLA4), along with lower T cell infiltration and higher macrophage infiltration. Genetic variation analyses revealed that the high-AMLPI samples associate with adverse variation events, including TP53 mutations, secondary NPM1 co-mutations, and copy number deletions. Biological interpretation indicated that ALDH2 and SPATS2L contribute significantly to AML patient survival, and their abnormal expression correlates with DNA methylation at cg12142865 and cg11912272. Drug response analyses revealed that different AMLPI samples tend to have different clinical selections, with low-AMLPI samples being more likely to benefit from immunotherapy. Finally, to facilitate broader access to our findings, a user-friendly and publicly accessible webserver was established and available at http://bioinfor.imu.edu.cn/amlpi. This server provides tools including TME-related AML driver genes mining, AMLPI construction, multi-dimensional validations, AML patients risk assessment, and figures drawing.
Identifying a small set of effective biomarkers from multi-omics data is important for the discrimination of different cell types and helpful for the early detection diagnosis of complex diseases. However, it is challenging to identify optimal biomarkers from the high throughput molecular data. Here, we present a method called protein-protein interaction affinity and co-expression network (PPIA-coExp), a linear programming model designed to discover context-specific biomarkers based on co-expressed networks and protein-protein interaction affinity (PPIA), which was used to estimate the concentrations of protein complexes based on the law of mass action. The performance of PPIA-coExp excelled over the traditional node-based approaches in both the small and large samples. We applied PPIA-coExp to human aging and Alzheimer's disease (AD) and discovered some important biomarkers. In addition, we performed the integrative analysis of transcriptome and epigenomic data, revealing the correlation between the changes in gene expression and different histone modification distributions in human aging and AD.
Background: Current identification of chronic myelogenous leukemia markers tends to mine diagnostic or prognostic biomarkers, ignoring susceptibility markers in normal samples. Objective: We aim to identify possible susceptibility markers for preventing chronic myelogenous leukemia. Methods: Functional links of H3K79me2 patterns and gene expression changes were inferred by correlation analyses. DNase-seq read distribution, transcription factor motifs, and their binding data were acquired via ceasBW and HOMER. Normalized transcription factor binding signals were submitted to a random forest algorithm to predict susceptibility gene expression changes. Three strategies were performed to validate the influence of low H3K79me2 signals on gene expression changes. Results: The gene-body H3K79me2 signals in normal samples were negatively related to gene expression changes during leukemogenesis (ρ=-0.92), regardless of gene lengths and expression levels. Characterization revealed that genes with lower H3K79me2 signals in normal samples have more open environments. Transcription factors GATA3, GATA4, TEAD1, TEAD3, TEAD4, and TRPS1 may induce the upregulation of up-susceptibility genes (ρ=0.95), and ASCL2, IRF4, IRF3, E2A, OCT4, and ZEB2 may mediate the downregulation of down-susceptibility genes (ρ=0.97). Enrichment analysis implied that the screened susceptibility genes were involved in leukemia-related pathways, and about 50% of leukemia stem cell differentially expressed genes were included in these genes. Besides, all hub genes extracted from susceptibility genes were well documented in different leukemia subtypes. Finally, the effect of H3K79me2 signals on gene expression changes were validated in a mouse model and three cell models. Conclusion: Low gene-body H3K79me2 signals in normal samples may serve as susceptibility markers for chronic myelogenous leukemia.
Lung adenocarcinoma is one of the deadliest tumors. Studies have shown that N6-methyladenosine RNA methylation regulators, as a dynamic chemical modification, affect the occurrence and development of lung adenocarcinoma. To investigate the relationship between mutations and expression levels of m6A regulators in lung adenocarcinoma, we investigated the mutations and expression levels of 38 m6A regulators. We found that mutations in m6A regulatory factors did not affect the changes in expression levels, and 19 differentially expressed genes were identified. All tumor samples were classified into two subtypes based on the expression levels of 19 differentially expressed m6A-regulated genes. Survival analysis showed significant differences in survival between the two subtypes. To explore the relationship between immune cell infiltration and survival in both subtypes, we calculated the infiltration of 23 immune cells in both subtypes, and we found that the subtype with high immune cell infiltration had better survival. We found that subtypes with low tumor purity and high stromal and immune scores had better survival. The m6A-related immune genes were identified by taking the intersection of differentially expressed genes and immune genes in the two isoforms and calculating the Pearson correlation coefficients between the intersecting immune genes and the differentially expressed m6A-regulated genes. Finally, a prognostic model associated with m6A and associated with immunity was developed using prognostic genes screened from m6A-associated immune genes. The predictive power of the model was evaluated and our model was able to achieve good prediction.
BACKGROUND:The accumulation of fatty acids in plants covers a wide range of functions in plant physiology and thereby affects adaptations and characteristics of species. As the famous woody oilseed crop, Acer truncatum accumulates unsaturated fatty acids and could serve as the model to understand the regulation and trait formation in oil-accumulation crops. Here, we performed Ribosome footprint profiling combing with a multi-omics strategy towards vital time points during seed development, and finally constructed systematic profiling from transcription to proteomes. Additionally, we characterized the small open reading frames (ORFs) and revealed that the translational efficiencies of focused genes were highly influenced by their sequence features.RESULTS:The comprehensive multi-omics analysis of lipid metabolism was conducted in A. truncatum. We applied the Ribo-seq and RNA-seq techniques, and the analyses of transcriptional and translational profiles of seeds collected at 85 and 115 DAF were compared. Key members of biosynthesis-related structural genes (LACS, FAD2, FAD3, and KCS) were characterized fully. More meaningfully, the regulators (MYB, ABI, bZIP, and Dof) were identified and revealed to affect lipid biosynthesis via post-translational regulations. The translational features results showed that translation efficiency tended to be lower for the genes with a translated uORF than for the genes with a non-translated uORF. They provide new insights into the global mechanisms underlying the developmental regulation of lipid metabolism.CONCLUSIONS:We performed Ribosome footprint profiling combing with a multi-omics strategy in A. truncatum seed development, which provides an example of the use of Ribosome footprint profiling in deciphering the complex regulation network and will be useful for elucidating the metabolism of A. truncatum seed oil and the regulatory mechanisms.
Abnormal histone modifications (HMs) can promote the occurrence of breast cancer. To elucidate the relationship between HMs and gene expression, we analyzed HM binding patterns and calculated their signal changes between breast tumor cells and normal cells. On this basis, the influences of HM signal changes on the expression changes of breast cancer-related genes were estimated by three different methods. The results showed that H3K79me2 and H3K36me3 may contribute more to gene expression changes. Subsequently, 2109 genes with differential H3K79me2 or H3K36me3 levels during cancerogenesis were identified by the Shannon entropy and submitted to perform functional enrichment analyses. Enrichment analyses displayed that these genes were involved in pathways in cancer, human papillomavirus infection, and viral carcinogenesis. Univariate Cox, LASSO, and multivariate Cox regression analyses were then adopted, and nine potential breast cancer-related driver genes were extracted from the genes with differential H3K79me2/H3K36me3 levels in the TCGA cohort. To facilitate the application, the expression levels of nine driver genes were transformed into a risk score model, and its robustness was tested via time-dependent receiver operating characteristic curves in the TCGA dataset and an independent GEO dataset. At last, the distribution levels of H3K79me2 and H3K36me3 in the nine driver genes were reanalyzed in the two cell lines and the regions with significant signal changes were located.
Acer palmatum (A. palmatum), a deciduous shrub or small arbour which belongs to Acer of Aceraceae, is an excellent greening species as well as a beautiful ornamental plant. In this study, a high-quality chromosome-level reference genome for A. palmatum was constructed using Oxford Nanopore sequencing and Hi-C technology. The assembly genome was ∼745.78 Mb long with a contig N50 length of 3.20 Mb, and 95.30 % (710.71 Mb) of the assembly was anchored into 13 pseudochromosomes. A total of 28,559 protein-coding genes were obtained, ∼90.02 % (25,710) of which could be functionally annotated. The genomic evolutionary analysis revealed that A. palmatum is most closely related to A. yangbiense and A. truncatum, and underwent only an ancient gamma whole-genome duplication event. Despite lacking a recent independent WGD, 25,795 (90.32 %) genes of A. palmatum were duplicated, and the unique/expanded gene families were linked with genes involved in plant-pathogen interaction and several metabolic pathways, which might underpin adaptability. A combined genomic, transcriptomic, and metabolomic analysis related to the biosynthesis of anthocyanin in leaves during the different season were characterized. The results indicate that the dark-purple colouration of the leaves in spring was caused by a high amount of anthocyanins, especially delphinidin and its derivatives; and the red colouration of the leaves in autumn by a high amount of cyanidin 3-O-glucoside. In conclusion, these valuable multi-omic resources offer important foundations to explore the molecular regulation mechanism in leaf colouration and also provide a platform for the scientific and efficient utilization of A. palmatum.
核小体是真核生物染色质的基本单位.核小体的精确定位影响了基因组序列对结合蛋白的可及性、转录、遗传复制和重组.了解核小体在基因组的准确位置对理解真核生物的生命活动过程有重要作用.本文基于核小体的序列和结构特征及统计物理理论,用统计物理模型预测了核小体的定位.利用统计物理和信息论原理计算了酿酒酵母(S.cerevisiae)、人类(H.sapiens)、秀丽隐杆线虫(C.elegans)和黑腹果蝇(D.melanogaster)数据集中序列片段的DNA局部结构的总能量,基于核小体序列与非核小体序列的总能量差异进行分类,通过10倍交叉验证进行了性能评估.结果显示该模型具有较好的识别效能.
Low temperature is one of the most prominent environmental factors affecting plant growth. As a deciduous arboreal tree, Acer palmatum has considerable ornamental and economic value; however, the molecular mechanisms underlying cold stress regulation in this species have yet to be determined. In this study, we performed Illumina high-throughput sequencing of 18 libraries obtained from A. palmatum subjected to cold treatment and subsequently identified a differentially expressed R2R3-MYB family gene, ApMYB77, which was then cloned and functionally characterized. The expression of ApMYB77 was induced by cold and drought treatments, and overexpression in A. thaliana enhanced the freezing tolerance (− 9 °C for 6 h) of transgenic plants compared with that of wild-type plants. The survival rates of transgenic plants (89% and 92%) were significantly higher than the wild-type plants (52%). Moreover, the transcript abundances of CBF-dependent regulatory pathway genes (AtCBF1, AtCBF2, AtCBF3, AtCBF4, AtCOR6.6A, AtCOR15B, AtCOR78, AtCOR414, and AtKIN1) were found to be significantly up-regulated in the transgenic lines. Enhanced abscisic acid (ABA)-dependent drought tolerance in transgenic plants was a further consequence of the overexpression of ApMYB77 following treatment of 15% polyethylene glycol for 4 d, compared with that in wild-type plants. In contrast, the expressions of genes (AtANAC072, AtDREB2A, AtERD1, AtMYB2, AtRD20, and AtRD29A) positively regulated by ABA were activated. Overall, the findings of this study indicate that ApMYB77 confers both freezing and drought tolerances.
Abnormal DNA methylation can alter the gene expression to promote or inhibit tumorigenesis in colon adenocarcinoma (COAD). However, the finding important genes and key sites of abnormal DNA methylation which result in the occurrence of COAD is still an eventful task. Here, we studied the effects of DNA methylation in the 12 types of genomic features on the changes of gene expression in COAD, the 10 important COAD-related genes and the key abnormal DNA methylation sites were identified. The effects of important genes on the prognosis were verified by survival analysis. Moreover, it was shown that the important genes were participated in cancer pathways and were hub genes in a co-expression network. Based on the DNA methylation levels in the ten sites, the least diversity increment algorithm for predicting tumor tissues and normal tissues in seventeen cancer types are proposed. The better results are obtained in jackknife test. For example, the predictive accuracies are 94.17 %, 91.28 %, 89.04 % and 88.89 %, respectively, for COAD, rectum adenocarcinoma, pancreatic adenocarcinoma and cholangiocarcinoma. Finally, by computing enrichment score of infiltrating immunocytes and the activity of immune pathways, we found that the genes are highly correlated with immune microenvironment.