Mallory-Denk bodies (MDBs) are protein aggregates commonly observed in chronic liver diseases, including liver fibrosis. However, the intrahepatic crosstalk driving MDB pathogenesis and fibrosis progression remains poorly understood. Using single-nucleus RNA sequencing (snRNA-seq), we identified significant cellular heterogeneity and a distinct hepatocyte subpopulation, termed MDB-associated hepatocytes (MAHs). MAHs were strongly correlated with hepatocellular carcinoma progression. Four hepatic stellate cell (HSC) subpopulations were defined, among which activated HSCs (aHSCs) represent a unique MDB-associated subtype. Moreover, we revealed a tightly connected axis involving MAHs, aHSCs, and Kupffer cells (KCs), which demonstrated that aberrant hepatocyte growth factor (HGF)/mesenchymal‒epithelial transition factor (MET) signaling contributes to MDB pathogenesis. Mechanistically, HGF secreted by aHSCs or KCs interacts with MET on ballooned MAHs and stimulates the HGF/MET downstream PI3K/AKT/NF-κB and STAT3 pathways via protein phosphorylation. The activated HGF/MET pathway promotes ubiquitin D (UbD) upregulation and the release of the proinflammatory cytokine TNFα which further promotes HGF transcription, establishing a positive feedback loop and contributing to MDB formation. Furthermore, aHSCs promote MDB pathogenesis by regulating STAT3 via the HGF/MET axis and increase HSC activation by stimulating TGFβ1 secretion, thereby accelerating fibrosis in 3D MDB organoid cultures. Notably, UbD deficiency (in UbD⁻/⁻ mice) suppressed HGF/MET signaling and MDB formation, leading to reduced liver fibrosis. Consistently, HGF/MET signaling was markedly elevated in human liver biopsies containing MDBs. Together, these findings provide unprecedented single-cell insights into liver cell reprogramming and intrahepatic crosstalk during MDB pathogenesis, and highlight the HGF/MET/UbD axis as a potential therapeutic target for chronic liver disease.
Unexplored biological matter-including uncharacterized genetic elements, molecular entities, and microbial components-remains poorly understood. Here, we use integrated multi-omics approaches to identify and characterize previously unrecognized protein products encoded by circular RNAs (circRNAs) in human tissue specimens and to delineate their roles in the progression of lung adenocarcinoma (LUAD). The transcription of precursor mRNA by RNA polymerase Ⅱ subunit A (RPB1) is crucial for the biogenesis of these potential circRNA-encoded proteins. Functional and translational analyses link their expression to distinct pathological stages of LUAD in patients. The protein RIPK1-98, encoded by circRIPK1, was identified as functionally distinct from its parental gene product, receptor-interacting serine/threonine kinase 1 (RIPK1). RIPK1-98 modulates cyclin-dependent kinase 2 (CDK2)-dependent cell-cycle regulation, thereby facilitating tumor proliferation in cellular and animal models. Together, these findings suggest that RIPK1-98 serves as a biomarker for cell-cycle progression in LUAD and highlight its potential as a therapeutic target to counteract resistance to first-line treatments, such as osimertinib.
Acute myeloid leukemia (AML) is a clinically aggressive hematologic malignancy driven by complex genetic and epigenetic aberrations. Circular RNAs (circRNAs), characterized by covalently closed structures and exceptional stability, have emerged as promising diagnostic biomarkers. However, existing circRNA-based predictive models largely depend on differential expression, overlooking the potential impact of higher-order chromatin organization on circRNA formation and function. Here, we propose a machine learning framework that integrates three-dimensional (3D) genome architecture to refine circRNA selection for AML prediction. By mapping 9,565 circRNAs onto a 3D chromatin model reconstructed from Hi-C data, we analyzed their spatial clustering and biological pathway enrichment. Eighteen pathways exhibited significant 3D aggregation of circRNAs, enabling radial stratification based on nuclear localization. Five circRNA panels were designed using complementary strategies combining expression, pathway, and spatial features. Cross-validation and external validation across six machine learning algorithms showed that the panel derived from the fifth radial layer (Panel-3DG-Radius5) achieved the most robust and consistent performance (ROC-AUC > 0.99). Integrating 3D genomic context reduced feature collinearity while enhancing biological interpretability. Overall, our study establishes a 3D genome-informed paradigm for circRNA biomarker discovery, demonstrating that spatial genome organization can substantially improve the precision and robustness of AML predictive modeling.
Cuproptosis is a copper-dependent regulated cell death pathway, but its connection to chromatin regulation is poorly understood. Here, we demonstrate that cuproptosis-sensitive tumor cells exhibit an "epigenetically primed" state with elevated chromatin accessibility and active histone marks. Multi-omics analyses reveal extensive chromatin reprogramming, including topologically associating domain (TAD) fusion and global reduction of enhancer-associated loops. We identify the chromatin remodeler CHD4, a core subunit of the NuRD complex, as a direct copper sensor. Copper ions bind to the CXXC domain of CHD4, triggering its ubiquitin-mediated degradation. As a negative regulator, CHD4 loss causes chromatin decompaction and de-represses the transcription factor HSF2, which directly transactivates the key cuproptosis executor FDX1. Genetic and pharmacological validations confirm the copper-CHD4-HSF2-FDX1 axis as a central regulator of cuproptosis susceptibility. In patient-derived models, high HSF2/FDX1 expression predicts enhanced response to cuproptosis inducers. Our work establishes an epigenetic mechanism linking copper sensing to cuproptosis and nominates the CHD4/HSF2/FDX1 axis as potential biomarkers and therapeutic targets for precision oncology.
Background: Molecular subtyping of breast cancer usually relies on transcriptomic profiles, a method constrained by limitations in robustness and clinical applicability. While somatic point mutations represent a stable genomic alternative, their predictive utility is hindered by high dimensionality, extreme sparsity, and weak single-gene associations. Methods: Here, we present deepGene-BC, a deep learning framework that synergizes a pathway-informed feature selection strategy with a hybrid neural network tailored for sparse binary data. To distill sparse genome-wide mutations into a compact and interpretable feature set, deepGene-BC integrates mutation recurrence filtering, curated pathway priors, and mutual information-based gene prioritization. These refined features are subsequently modeled using a specialized hybrid architecture designed to capture complex linear effects, feature interactions, and higher-order nonlinear patterns. Results: When benchmarked against an independent test set (n = 273) from the TCGA breast cancer cohort, deepGene-BC achieved an overall accuracy of 77.3% and an average sensitivity of 75.2%, accompanied by a strong overall discriminative performance (macro-averaged AU-ROC = 0.94, 95% CI: 0.92-0.96). Conclusions: By effectively combining biologically informed feature engineering with deep learning, deepGene-BC holds significant promise for non-invasive molecular stratification and precision oncology.
Gestational diabetes mellitus (GDM) remains a prevalent and heterogeneous pregnancy complication with limited strategies for early identification. We aimed to investigate efficient approaches for early prediction of GDM with clinical and genetic risk factors. A previously developed machine-learning model based on clinical characteristics achieved an area under the curve (AUC) of 0.77. To improve predictive accuracy, we further collected non-invasive prenatal testing (NIPT) results from 595 pregnant women (295 with GDM, 300 without). A cumulative polygenic risk score (PRS) was calculated using 1,170 selected single nucleotide variants (SNVs). Logistic regression, support vector machines, random forest, decision tree, linear model and naïve Bayes machine learning models were employed. External validation was performed with an additional 2,350 blood samples independently collected from two other centers. Logistic regression analysis showed that the PRS alone achieved an AUC of 0.75 for GDM discrimination. From cell-free DNA (cfDNA) sequencing performed during NIPT, we identified 357 gene transcripts with differential coverage at transcription start sites. A cfDNA-based linear model achieved an AUC of 0.83 using a subset of 166 signature genes, which reached 0.85 when combined with clinical features. Integration of clinical features, cfDNA, and SNVs yielded the highest performance using a random forest model (AUC = 0.89, specificity = 0.74, sensitivity = 0.89). For external validation, a clinically practical model incorporating clinical features and cfDNA achieved an AUC of 0.83 using linear approach. Our GDM prediction model has reached high accuracy fully using accessible clinical and genetic data routinely generated from current antenatal testing, enabling early screening and interventions for women at risk.
Genetic/genomic manipulation techniques (gene transfer/delivery, gene editing, etc.) have become more and more mature, and the illegal use as gene doping in sports has drawn attentions. World Anti-Doping Agency (WADA) strictly prohibits gene doping, and has issued guideline on quantitative real-time PCR (qPCR) detections. However, the technical feature of qPCR makes it difficult to detect new doping targets, and codon changes on targets may also affect detection efficiency. Here, we prepare standard materials for genomic and transgenic versions of human EPO (hEPO) gene, and design qPCR primers to check the consequences of codon changes on gene doping detection. We confirm that carefully designed qPCR assays could indeed capture transgene signal, but codon changes on the transgene could severely undermine detection efficiency. We have also mimicked real world gene doping scenario by mixing genomic and transgenic versions of hEPO, and qPCR could detect wild-type but not codon-changed transgenes. As a method validation for such a challenge, we also use Sanger sequencing to confirm that sequencing could easily capture gene doping even for codon-changed transgenes. Our study confirms that codon changes will challenge qPCR-based gene doping detection, and calls for un-biased detection tools based on high-throughput sequencing in the future.
AIMS:The challenges in identifying functional variants from genome-wide association studies (GWAS) and unraveling regulatory mechanisms in schizophrenia research persist, particularly in intronic regions. A non-coding regulatory variant, rs1399178, associated with schizophrenia risk, is identified and its impact on NRF1 binding is investigated. METHODS:This study focuses on schizophrenia GWAS risk loci, using functional genomics, expression analyses and structural analysis to identify 736 schizophrenia risk single-nucleotide polymorphisms (SNPs) that disrupt transcription factor (TF) binding. RESULTS:Among these SNPs, rs1399178 stands out as a bifunctional intergenic SNP that can switch between acting as a promoter and an enhancer, potentially influencing MLH1 and LRRFIP2 expression via expression quantitative trait loci analysis. Importantly, mutation of the G allele of rs1399178 to A significantly diminishes its binding affinity to nuclear respiratory factor 1 (NRF1). Structural analysis provides further insight into this alteration in the protein-nucleic acid complex formation. CONCLUSION:Based on our data, a model is proposed in which rs1399178 confers schizophrenia risk by modifying NRF1 binding profiles, thereby regulating the abundance of target genes through promoter-enhancer switching. This study provides novel insights into the regulatory mechanisms of schizophrenia risk variants, highlighting the intricate nature of genetic interactions and potential therapeutic targets.
Phase separation (PS) is essential in cellular processes and disease mechanisms, highlighting the need for predictive algorithms to analyze uncharacterized sequences and accelerate experimental validation. Current high-accuracy methods often rely on extensive annotations or handcrafted features, limiting their generalizability to sequences lacking such annotations and making it difficult to identify key protein regions involved in PS. We introduce Phase Separation’s Transfer-learning Prediction (PSTP), which combines conformational embeddings with large language model embeddings, enabling state-of-the-art PS predictions from protein sequences alone. PSTP performs well across various prediction scenarios and shows potential for predicting novel-designed artificial proteins. Additionally, PSTP provides residue-level predictions that are highly correlated with experimentally validated PS regions. By analyzing 160 000+ variants, PSTP characterizes the strong link between the incidence of pathogenic variants and residue-level PS propensities in unconserved intrinsically disordered regions, offering insights into underexplored mutation effects. PSTP’s sliding-window optimization reduces its memory usage to a few hundred megabytes, facilitating rapid execution on typical CPUs and GPUs. Offered via both a web server and an installable Python package, PSTP provides a versatile tool for decoding protein PS behavior and supporting disease-focused research.
Since the early 20th century, the concept of doping was first introduced. To achieve better athletic performance, chemical substances were used. By the mid-20th century, it became gradually recognized that the illegal use of doping substances can seriously endangered athletes' health and compromised the fairness of sports competitions. Over the past 30 years, the World Anti-Doping Agency (WADA) has established corresponding rules and regulations to prohibit athletes from using doping substances or restrict the use of certain drugs, and isotope, chromatography, and mass spectrometry techniques were accredited to detect doping substances. With the development of gene editing technology, many genetic diseases have been effectively treated, but enabled by the same technology, doping has also the potential to pose a threat to sports in the form of gene doping. WADA has explicitly indicated gene doping in the Prohibited List as a prohibited method (M3) and approved qPCR detection. However, gene doping can easily evade detection, if the target genes' upstream regulatory elements are considered, the task became more challenging. Hi-C experiment driven 3D genome technology, through perspectives such as topologically associating domain (TAD) and chromatin loop, provides a more comprehensive and in-depth understanding of gene regulation and expression, thereby better preventing the potential use of 3D genome level gene doping. In this work, we will explore gene doping from a different perspective by analyzing recent studies on gene doping and explore related genes under 3D genome.
Due to the lack of validated universal seizure markers, population-level prediction methods often exhibit limited performance. This study proposes homologous microstate dynamic attributes as a generalized, subject-independent seizure marker. Homologous microstate dynamic attributes were extracted using a novel spatiotemporal graph convolutional network (ST-GCN) model for subject-independent seizure prediction. An online deployment stage was introduced to validate the model's clinical applicability. The online deployment stage demonstrated that the model achieved sensitivities of 96.79% and 98.84% on the private dataset and Siena dataset, respectively. The ST-GCN model successfully predicts seizures in a subject-independent manner, demonstrating its potential as a generalized tool for seizure prediction in clinical settings. This study indicates that dynamics within homologous microstates can serve as a universal predictive biomarker for seizures, expanding microstate research beyond transition patterns. It also provides a practical template for clinical seizure prediction models.
Accurate cancer survival prediction is crucial in devising optimal treatment plans and offering individualized care to improve clinical outcomes. Recent researches confirm that integrating heterogenous cancer data such as histopathological images and genomic data, can enhance our understanding of cancer progression and provides a multimodal perspective on patient survival chances. However, existing methods often over-look the fundamental aspects of multimodal data, i.e., consistency and complementarity, which in consequence significantly hinder advancements in cancer survival prediction. To address this issue, we represent DRLSurv, a novel multimodal deep learning method that leverages disentangled representation learning for precise cancer survival prediction. Through dedicated deep encoding networks, DRLSurv decomposes each modality into modality-invariant and modality-specific representations, which are mapped to common and unique feature subspaces for simultaneously mining the distinct aspects of cancer multimodal data. Moreover, our method innovatively introduces a subspace-based proximity contrastive loss and re-disentanglement loss, thus ensuring the successful decomposition of consistent and complementary information while maintaining the multimodal fidelity during the learning of disentangled representations. Both quantitative analyses and visual assessments on different datasets validate the superiority of DRLSurv over existing survival prediction approaches, demonstrating its powerful capability to exploit enriched survival-related information from cancer multimodal data. Therefore, DRLSurv not only offers a unified and comprehensive deep learning framework for advancing multimodal survival predictions, but also provides valuable insights for cancer prognosis and survival analysis.
Aging is a complex biological process characterized by increased inflammation and susceptibility to various age-related diseases, including cognitive decline, osteoporosis, and type 2 diabetes. Exercise has been shown to modulate mitochondrial function, immune responses, and inflammatory pathways, thereby attenuating aging through the regulation of exerkines secreted by diverse tissues and organs. These bioactive molecules, which include hepatokines, myokines, adipokines, osteokines, and neurokines, act both locally and systemically to exert protective effects against the detrimental aspects of aging. This review provides a comprehensive summary of different forms of exercise for older adults and the multifaceted role of exercise in anti-aging, focusing on the biological functions and sources of these exerkines. We further explore how exerkines combat aging-related diseases, such as type 2 diabetes and osteoporosis. By stimulating the secretion of these exerkines, exercise supports healthy longevity by promoting tissue homeostasis and metabolic balance. Additionally, the integration of exercise-induced exerkines into therapeutic strategies represents a promising approach to mitigating age-related pathologies at the molecular level. As our understanding deepens, it may pave the way for personalized interventions leveraging physical activity to enhance healthspan and improve quality of life.
Schizophrenia is a polygenic complex disease with a heritability as high as 80 %, yet the mechanism of polygenic interaction in its pathogenesis remains unclear. Studying the interaction and regulation of schizophrenia susceptibility genes is crucial for unraveling the pathogenesis of schizophrenia and developing antipsychotic drugs. Therefore, we developed a bioinformatics method named GRACI (Gene Regulation Analysis based on Causal Inference) based on the principles of information theory, a causal inference model, and high order chromatin 3D conformation. GRACI captures the interaction and regulatory relationships between schizophrenia susceptibility genes by analyzing genotyping data. Two datasets, comprising 1459 and 2065 samples respectively, were analyzed, and the gene networks from both datasets were constructed. GRACI showcased superior accuracy when compared to widely adopted methods for detecting gene-gene interactions and intergenic regulation. This alignment was further substantiated by its correlation with chromatin high-order conformation patterns. Using GRACI, we identified three potential genes - KCNN3 , KCNH1 , and KCND3 - that are directly associated with schizophrenia pathogenesis. Furthermore, the results of GRACI on the standalone dataset illustrated the method ' s applicability to other complex diseases. GRACI download: https://github.com/liuliangjie19/GRACI
Non-alcoholic fatty liver disease(NAFLD)is a liver condition that is widely prevalent across the world.A considerable number of people with NAFLD have the potential to prog-ress to a more severe form of the condition known as nonalcoholic steatohepatitis(NASH),accompanied by bridging fibrosis.This advancement is more likely if the patient has metabolic risk factors such as obesity or type 2 diabetes that deteriorate over time.Additionally,even slight inflammation or fibrosis in NAFLD can significantly increase the likelihood of progression compared to steatosis alone.This underscores the importance of revising the present methods of monitoring NAFLD patients to ensure early detection and effective management of the disease.
Reliable and ultra-fast DNA and RNA sequencing have been achieved with the emergence of high-throughput sequencing technology. When combining the results of DNA and RNA sequencing for tumor cells of cancer patients, neoantigens that potentially stimulate the immune response of either CD4+ or CD8+ T cells can be identified. However, due to the abundance of somatic mutations and the high polymorphic nature of human leukocyte antigen (HLA) it is challenging to accurately predict the neoantigens. Moreover, comparing to HLA-I presented peptides, the HLA-II presented peptides are more variable in length, making the prediction of HLA-II loaded neoantigens even harder. A number of computational approaches have been proposed to address this issue but none of them considers the DNA origin of the neoantigens from the perspective of 3D genome. Here we investigate the DNA origins of the immune-positive and non-negative HLA-II neoantigens in the context of 3D genome and discovered that the chromatin 3D architecture plays an important role in more effective HLA-II neoantigen prediction. We believe that the 3D genome information will help to increase the precision of HLA-II neoantigen discovery and eventually benefit precision and personalized medicine in cancer immunotherapy.
Background The impact of the gut microbiome on the initiation and intensity of immune-related adverse events (irAEs) prompted by immune checkpoint inhibitors (ICIs) is widely acknowledged. Nevertheless, there is inconsistency in the gut microbial associations with irAEs reported across various studies. Methods We performed a comprehensive analysis leveraging a dataset that included published microbiome data ( n = 317) and in-house generated data from 16S rRNA and shotgun metagenome samples of irAEs ( n = 115). We utilized a machine learning-based approach, specifically the Random Forest (RF) algorithm, to construct a microbiome-based classifier capable of distinguishing between non-irAEs and irAEs. Additionally, we conducted a comprehensive analysis, integrating transcriptome and metagenome profiling, to explore potential underlying mechanisms. Results We identified specific microbial species capable of distinguishing between patients experiencing irAEs and non-irAEs. The RF classifier, developed using 14 microbial features, demonstrated robust discriminatory power between non-irAEs and irAEs (AUC = 0.88). Moreover, the predictive score from our classifier exhibited significant discriminative capability for identifying non-irAEs in two independent cohorts. Our functional analysis revealed that the altered microbiome in non-irAEs was characterized by an increased menaquinone biosynthesis, accompanied by elevated expression of rate-limiting enzymes menH and menC . Targeted metabolomics analysis further highlighted a notably higher abundance of menaquinone in the serum of patients who did not develop irAEs compared to the irAEs group. Conclusions Our study underscores the potential of microbial biomarkers for predicting the onset of irAEs and highlights menaquinone, a metabolite derived from the microbiome community, as a possible selective therapeutic agent for modulating the occurrence of irAEs.
Phase separation (PS) is essential in various biological processes, necessitating high-accuracy predictive algorithms to study numerous uncharacterized sequences and accelerate experimental validation. However, many recent prediction methods face challenges in generalizability due to their reliance on engineered features, and accurately identifying protein regions involved in PS remains difficult. We propose PSTP, which employs a dual-language model embedding strategy and a lightweight attention model. PSTP achieves identification of 84% of PS regions in PhaSePro and demonstrates ~50% improvement in correlation coefficient over existing models. It shows robust performance across different types of PS proteins, and shows the potential for predicting artificial proteins. By analyzing 160,000 variants, PSTP characterizes the link between the incidence of pathogenic variants and residue-level PS propensities. PSTP's predictive power and broad applicability make it a valuable tool for accelerating our understanding of biomolecular condensates, and their mechanisms underlying disease development. ### Competing Interest Statement The authors have declared no competing interest.
Recent progress in gene editing has enabled development of gene therapies for many genetic diseases, but also made gene doping an emerging risk in sports and competitions. By delivery of exogenous transgenes into human body, gene doping not only challenges competition fairness but also places health risk on athletes. World Anti-Doping Agency (WADA) has clearly inhibited the use of gene and cell doping in sports, and many techniques have been developed for gene doping detection. In this review, we will summarize the main tools for gene doping detection at present, highlight the main challenges for current tools, and elaborate future utilizations of high-throughput sequencing for unbiased, sensitive, economic and large-scale gene doping detections. Quantitative real-time PCR assays are the widely used detection methods at present, which are useful for detection of known targets but are vulnerable to codon optimization at exon-exon junction sites of the transgenes. High-throughput sequencing has become a powerful tool for various applications in life and health research, and the era of genomics has made it possible for sensitive and large-scale gene doping detections. Non-biased genomic profiling could efficiently detect new doping targets, and low-input genomics amplification and long-read third-generation sequencing also have application potentials for more efficient and straightforward gene doping detection. By closely monitoring scientific advancements in gene editing and sport genetics, high-throughput sequencing could play a more and more important role in gene detection and hopefully contribute to doping-free sports in the future.