Drug-Target Interaction (DTI) is a crucial aspect of pharmaceutical development. However, biochemical experiments are prohibitively expensive to identify these interactions on a large scale, while the computational approach is still on the way to making a highly reliable prediction. For the purpose of promoting prediction accuracy, drug-related molecular networks are gradually introduced to this task to furnish valuable information. We hypothesized that integrating structural and systemic biological attributes could effectively enhance the performance of DTI prediction and proposed a novel DTI prediction model, SSGraphDTI, which integrated two aforementioned attributes. Specifically, the structural attributes of drugs and targets are extracted using independent convolutional neural network based models from the Simplified Molecular Input Line Entry System of drugs and the amino acid sequences of targets, respectively. Meanwhile, the systemic biological attributes of drug-target pairs are obtained through graph representation learning on the dynamically constructed heterogeneous drug-target interaction network. SSGraphDTI was meticulously trained and rigorously tested on the benchmark Dataset_DrugBank, achieving an improvement of approximately 1.0% across five metrics compared to recent comparable methods. These results underscore the potential of combining both structural and systemic information for accurate DTI prediction. Benefiting from the fact that the input consists solely of structural data without requiring interaction information, the model effectively addresses the "cold-start problem" in drug discovery. Furthermore, by extracting systemic attributes directly from the dynamically constructed DTI networks, the model maintains strong predictive performance even when data is limited.
Abstract Protein S-palmitoylation is a reversible lipid modification that regulates protein localization, trafficking, and signaling. Its dysregulation has been implicated in cancer and therapeutic resistance, making accurate site annotation important for understanding disease-related regulatory mechanisms. However, experimental identification of S-palmitoylation sites remains labor-intensive, highlighting the need for computational tools that can support large-scale candidate-site prioritization. S-palmitoylation-site recognition requires the integration of multiple levels of information surrounding candidate cysteine residues, including sequence context, local protein properties and structural features. Here, we present Deep-Palm, a deep learning framework that integrates four complementary branches: amino acid sequence, physicochemical properties, ESM embedding and spatial structure. On the testing set, Deep-Palm achieved an AUC of 0.950 and outperformed the compared predictors, including pCysMod, MusiteDeep and GPS-Palm. Deep-Palm also showed stable performance across diverse cysteine sequence contexts and Gene Ontology functional groups. Feature-level analyses revealed distinct spatial structural and protein property patterns between palmitoylated and non-palmitoylated peptide windows. Independent mass spectrometry datasets further demonstrated the ability of Deep-Palm to sensitively identify previously unannotated S-palmitoylation sites. Together, Deep-Palm provides an accurate and biologically informative framework for S-palmitoylation-site prediction, facilitating the discovery of novel candidate S-palmitoylation sites.
Osteoarthritis (OA) is a complex degenerative joint disease for which early diagnosis and clear molecular characterization remain limited. DNA methylation has been increasingly recognized as an important regulatory factor in OA pathogenesis. In this study, we proposed an integrative computational framework combining statistical analysis, machine learning, deep learning, and functional genomics to identify and validate OA-associated genes and methylation biomarkers for diagnostic and biological interpretation. Candidate CpG sites were obtained using two complementary strategies: differential methylation analysis and selection of loci located near transcription start sites of previously reported OA-related genes. Key features were further refined using support vector machine recursive feature elimination and random forest algorithms. Based on the selected loci, we developed a feature-fusion diagnostic model that combines Transformer and convolutional neural networks with adaptive weighting to capture both global dependency structures and local methylation patterns. A panel of 220 methylation sites demonstrated stable and reproducible diagnostic performance in an independent cohort. Functional annotation and pathway analysis highlighted several established OA-associated genes, including TGFBR2, SMAD3, PPARG, and MAPK3, and suggested INHBB as a potential novel effector gene, with additional support for AMH and INHBE involvement. Overall, this study presents a robust methylation-based framework for identifying key OA-associated genes and provides new insights into the epigenetic mechanisms underlying OA.
Automatic classification and segmentation of lesions in colonoscopy images is an important research direction in computer-aided diagnosis and plays a crucial role in clinical applications. However, most existing methods fail to sufficiently explore the relationship between lesion types and shapes, and the boundaries of lesions are often unclear. To address these problems, this paper proposes a multi-task collaborative network—MTCNet. MTCNet optimizes both classification task and segmentation task, and enhances the boundary awareness in the segmentation branch. To better exploit the relationship between boundary shapes and lesion categories, We propose a deep cooperation strategy based on Multi-Scale Transformer module. Through this module, the segmentation and classification tasks can be jointly optimized, enabling effective interaction between the two tasks. Meanwhile, to further enhance the model’s ability to recognize lesion boundaries, an Edge-Guided Attention (EGA) module based on the Laplacian algorithm is proposed. In addition, we construct a multi-center dataset PADset. Experimental results demonstrate that the proposed method outperforms commonly used approaches in both segmentation and classification tasks on the PADset dataset and the public SUN dataset.
Drug-Target Interactions (DTIs) identification is a crucial phase in drug development to screen potential drugs and targets and improve the efficiency of drug development. Accompanied by advances in Artificial Intelligence (AI) technology, Deep Learning (DL) methods have been introduced to the field and designed for DTI prediction, which can expedite the Research and Development (R&D) cycle. However, with the diverse DTIs involving various types of drugs and target protein molecules, it is still challenging to reveal the rules behind those interactions and further make reliable suggestions for the pharmaceutical industry. Currently, AI technology is expected to promote the ability to characterize the molecules against limited knowledge about DTIs. In this study, we accordingly propose a novel DTI prediction method named TopoPharmDTI, dedicated to better characterizing both the drug and target molecules, where an innovative dual-tier drug features fusion strategy is designed for drug molecules and a most adaptive Large Language Model (LLM) for target protein sequence representation is chosen under the evaluation of six top models by far. Our method achieves considerable prediction performance and a promising capability to identify the binding domains on the target protein. The method is freely available at http://ex.nenucompbio.com/TopoPharmDTI.html.
Central nervous system (CNS) drug discovery is constrained by an immense and sparse chemical search space. Meanwhile, molecules that simultaneously achieve brain penetration, target efficacy, and synthesizability are extremely scarce. However, existing generative models rarely couple rigorous multi-objective control with robustness to distribution shift, limiting their reliability in realistic CNS design. We introduce D²G-TO, a task-aware and out of-distribution (OOD)-guided discrete graph diffusion framework that unifies multi-pharmacological properties with structural distribution guidance. A novel Structural Similarity Guidance mechanism steers generation toward in-distribution regions while repelling OOD modes, maintaining structural distributional consistency in realistic scenarios. Across BBBP, BACE, and QM9 benchmarks, D²G-TO achieves strong validity, diversity, and other metrics. In an Alzheimer's disease case study, we subject the generated molecules to cross-property pharmacological prediction and systematic ADMET profiling, followed by structure based molecular docking against BACE-1 to assess binding-mode plausibility. D²G-TO identifies candidates that jointly satisfy blood–brain barrier permeability, β-site amyloid precursor protein cleaving enzyme 1 inhibition, and synthetic accessibility. Thus, D²G-TO has the potential to serve as an efficient in silico engine for early-stage CNS drug design. The code is available at https://github.com/zhaix922/DDG_TO.
Aneuploidy is intrinsically associated with nascent polyploidy, yet its evolutionary relevance in this specific context remains unknown. We assessed evolvability, defined as “progeny fitness,” of transgenerationally selected aneuploid and euploid cohorts in a synthetic allotetraploid wheat under both regular condition and three abiotic stresses. We find that ~30% of the aneuploid cohorts manifested significantly higher evolvability than euploidy even under regular normal condition, and the vast majority of individuals that survived any of the three stresses were aneuploidies with loss/gain of specific chromosomes; the few euploid survivors were all progenies of stress-tolerant parental aneuploidies, which contained the same inherited multiple structural variants. Physiological assay revealed stress-induced shifting of cellular physiological profiles which move toward or transgress those of the more tolerant diploid parent in the stress-tolerant cohorts. Whole-genome resequencing ruled out nucleotide-level mutation as an underpinning. This study demonstrates that aneuploidy may serve as a transit springboard contributing to the establishment of nascent polyploidy by imparting its stress tolerance and higher evolvability properties to euploid progenies.
Alzheimer’s disease (AD) is a progressive and irreversible neurodegenerative disease that has become one of the most severe and costly global public health crises today. Accurate diagnosis of AD progression from neuroimaging remain formidable challenge, particularly in multi-class scenarios requiring the delineation of subtle stage transitions. In this work, we propose a dynamic cross-modal fusion (DCMF) framework for multi-stage diagnosis of AD. The framework introduces three primary innovations: (1) adaptive single-modality representation learning, which enhances the extraction of disease-specific features from specific modalities; (2) dynamic, stage-aware cross-modal fusion, which captures synergistic interactions and adaptively aggregates complementary information; and (3) a joint supervision and feature alignment strategy designed to mitigate gradient vanishing and enforce the learning of highly discriminative representations. We conducted extensive evaluations to assess the performance of the proposed framework. Experimental results demonstrate that our method consistently outperforms state-of-the-art baselines, surpassing the second-best mode by 3.6 https://github.com/NENUBioCompute/DCMF .
MOTIVATION:Current epigenetic clocks face a trade-off between predictive accuracy and biological interpretability, often relying on dataset-specific correction to generalize across cohorts. We propose GT-Mamba, a novel architecture that integrates a Structure-Aware Graph Transformer with the Mamba state space model. This design captures CpG topological correlations and genome-wide long-range dependencies. RESULTS:GT-Mamba demonstrates strong out-of-the-box robustness across heterogeneous independent validation cohorts, achieving a weighted average MAE of 4.43 years. Notably, it effectively generalizes to EPIC 850k arrays despite partial feature missingness, and maintains consistent performance across homologous age distribution shifts (MAE 2.94 years in a young cohort). Ablation studies confirm that graph topology contributes to improved robustness against noise. Mechanistic analysis suggests that the model captures methylation patterns associated with both developmental and functional processes. AVAILABILITY:Source code and pre-trained models are freely available at https://github.com/NENUBioCompute/GT-Mamba and archived on Zenodo (DOI: 10.5281/zenodo.19703155).
Freezing of gait (FoG) in Parkinson's disease is a brief but hazardous gait failure that often precedes falls. For wearable cueing or other closed-loop assistance, a detector that reacts only after FoG onset is usually too late; the more useful task is to recognize the pre-freezing transition from physiological signals. This study presents PreFoGNet, a dual time-frequency deep learning framework for early FoG prediction using plantar pressure signals. The temporal stream combines a multi-scale Inception encoder with a bidirectional Mamba module to capture both short contact-related transients and several-second gait deterioration without the quadratic cost of attention. In parallel, the frequency stream uses band-wise spectral modeling and attention-based gating to emphasize physiologically meaningful changes in the locomotion, freeze-related, and high-frequency bands. On the WearGait-PD dataset, with a 2 s prediction horizon and subject-wise evaluation, PreFoGNet achieved a sensitivity of 93.94%, a specificity of 89.76%, a G-Mean of 0.9183, and an AUC-ROC of 0.9607. It outperformed classical machine-learning and deep learning baselines, and retained usable performance under moderate noise and single-channel loss. Additional horizon analysis showed that plantar pressure contains a stable pre-freezing signature within 0-3 s before onset, with a practical prediction boundary of approximately 6-7 s. These findings suggest that time-frequency modeling of plantar pressure is a promising signal-processing route for wearable FoG early-warning systems. ### Competing Interest Statement The authors have declared no competing interest. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The WearGait-PD dataset is available on Synapse (SAGE Bionetworks) at the following URL: https://www.synapse.org/Synapse:syn52540892/wiki/623751. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present work are contained in the manuscript. The WearGait-PD dataset is available on Synapse (SAGE Bionetworks) at the following URL: https://www.synapse.org/Synapse:syn52540892/wiki/623751. Natural Science Foundation of China, 82571771 Natural Science Foundation of Shanghai, 25ZR1401167
Meiotic crossover (CO) exchanges genetic information between homologs, thereby promoting genetic diversity among offspring. COs are non-randomly distributed across chromosomes, tending to occur in euchromatin, but rarely in heterochromatin. In plants, H3 lysine 27 monomethylation (H3K27me1) is crucial for maintaining heterochromatin condensation and genome stability in somatic cells; however, its role in germline cells remains to be determined. Here, we demonstrate that the plant-specific H3K27 mono-methyltransferases ATXR5/6 (ARABIDOPSIS TRITHORAX-RELATED PROTEIN 5/6) play an important role in inhibiting CO formation in meiotic heterochromatin. In atxr5 atxr6, both ZMM-dependent Type I COs and ZMM-independent Type II COs are significantly increased. We further observed decondensation, decreased H3K27me1 signals, and specifically compromised non-CG methylation in atxr5 atxr6 meiotic heterochromatin. Unexpectedly, in contrast to their roles in somatic cells, where ATXR5/6 primarily regulate heterochromatin condensation and gene silencing without influencing DNA methylation, in meiocytes, ATXR5/6 mainly function in suppressing recombination and preserving heterochromatic DNA methylation without directly regulating gene expression. Moreover, loss of Type II CO regulator MMS AND UV SENSITIVE 81 (MUS81) leads to pericentromeric fragmentation and polyad formation during meiosis in the absence of ATXR5/6, indicating that MUS81 is critical for resolving atypical recombination intermediates in pericentromeric heterochromatin. Taken together, our results provide insights into the roles of ATXR5/6 in repressing meiotic recombination within heterochromatin by regulating chromosome compaction and modifications.
The rational design of small molecules is central to drug discovery, yet current artificial intelligence (AI) methodologies for generating three-dimensional (3D) molecules are often siloed, focusing on either de novo design or fragment-based design. The lack of a holistic framework limits AI’s application across the complex and multi-step pipeline spanning from novel scaffold identification to lead compound optimization, and prevents AI from effectively learning from the entire process. Here, we introduce UniLingo3DMol, a language model for 3D molecular generation, empowered by fragment permutation-capable molecular representation alongside multi-stage and multi-task training strategy. This integrated design enables UniLingo3DMol to seamlessly span both de novo and fragment-retained molecular design, demonstrating superior performance over existing generation models in in silico evaluations across more than 100 diverse biological targets. We further leveraged UniLingo3DMol in the design of inhibitors targeting CBL-B, a crucial immune E3 ubiquitin ligase and attractive immunotherapy target. This strategy led to a lead compound demonstrating excellent in vitro activity and robust in vivo anti-tumor efficacy. Our findings establish UniLingo3DMol as a generalized and powerful platform, showing the strong potential to advance AI-driven drug discovery. ### Competing Interest Statement The authors have declared no competing interest. Beijing Municipal Science and Technology Commission, Z241100007724005
Accurately predicting drug-target interactions (DTI) is crucial for drug discovery and can reduce drug development costs. Recent deep learning-based DTI predictions have demonstrated promising performance, but they still face two challenges: (i) The over-reliance on the extraction of local features and insufficient learning of global features limit the model's performance. (ii) The lack of effective fusion of drug-target interaction features leads to the lack of interpretability of the model. To address these challenges, we propose a new model for predicting drug-target interactions based on multi-order gated convolution and multi-attention fusion, MGMA-DTI. The drug feature encoder obtains a two-dimensional molecular graph based on the drug's SMILES string and uses a graph convolutional neural network to encode the drug features. The protein encoder is based on a multi-order gated convolution, which enhances the model's ability to capture global feature between amino acid sequences. In order to better achieve interactive learning between drugs and proteins, we designed a multi-attention fusion module that effectively captures the drug-target interaction features. Experimental results show that MGMA-DTI outperforms other baseline models on three benchmark datasets: BindingDB, BioSNAP, and Human. Case studies further demonstrate that the model provides valuable insights for drug discovery. In addition, our model provides molecular-level interpretability, which can provide more scientifically meaningful guidance.
DNA methylation serves as a pivotal epigenetic mechanism that dynamically regulates gene expression, thereby influencing development, aging, and the onset of various diseases. In recent years, epigenetic clocks based on DNA methylation patterns have emerged as powerful tools for estimating biological age. However, many existing models still face limitations in predictive accuracy and offer limited insight into the molecular underpinnings of age-related disorders. To address these challenges, we developed a high-precision methylation-based aging clock using genome-wide DNA methylation data from the Illumina 450K array, encompassing a wide age range and multiple tissue types. Leveraging a deep learning framework based on the Mamba architecture, we constructed an age prediction model that achieves state-of-the-art performance in biological age estimation. In addition to enhanced predictive accuracy, the model offers a valuable framework for evaluating the efficacy of anti-aging interventions and exploring the epigenetic mechanisms that underlie aging and disease progression.
The tissue specificity of DNA methylation refers to the significant differences in DNA methylation patterns in different tissues. This specificity regulates gene expression, thereby supporting the specific functions of each tissue and the maintenance of normal physiological activities. Abnormal tissue-specific patterns of DNA methylation are closely related to age-related diseases. This abnormal methylation pattern affects the regulation of gene expression, which may lead to changes in cell function and promote the occurrence of pathological conditions. By analyzing the differences in these methylation patterns, key CpG sites for disease diagnosis can be effectively screened. The main goal of this paper is to use the characteristics associated with tissue-specific abnormal expression and disease to construct an age-related disease diagnosis model. First, we combined chi-square tests and logistic regression to identify tissue-specific and disease-specific CpG sites, laying the foundation for accurate medical diagnosis, and verified the biological relevance of these CpG sites through enrichment analysis. Then we used the Transformer model to fit these CpG sites and realized the automatic diagnosis of age-related diseases. Our work proves that the tissue specificity of DNA methylation has the potential to diagnose age-related diseases, and proves the scientific nature of our proposed diagnostic method from a biological perspective.
Accurately predicting the impact of point mutations on protein thermodynamic stability is essential for understanding structure-function relationships and guiding protein design. This challenge is particularly acute for transmembrane proteins (TMPs), which play vital roles in cellular signaling and drug targeting but remain underrepresented in structural databases. Existing predictors often rely on three-dimensional structures or multiple sequence alignments, limiting their applicability to TMPs due to poor structural coverage and alignment quality. Here, we present MEMO-Stab2, a fast and structure-independent deep learning framework for predicting mutation-induced stability changes in TMPs. MEMO-Stab2 reformulates the task as a binary classification problem, distinguishing destabilizing from neutral mutations based on a ΔΔG threshold of 0.4 kcal/mol. The model integrates multiview features within a Transformer-based architecture, utilizing embeddings from multiple pretrained protein language models (PLMs) and PLM-based structural predictions. By leveraging PLMs, it operates without requiring experimental 3D structures or explicit multiple sequence alignments, implicitly capturing both evolutionary and structural contexts from the amino acid sequence alone. Across internal and external transmembrane mutation data sets, MEMO-Stab2 consistently outperforms existing tools, including specialized predictors and a state-of-the-art general model even after it was fine-tuned on the same domain-specific data, achieving an F1 score of 0.92 on an internal benchmark. Further analyses confirm the model's robustness and specificity. It demonstrates strong generalization across diverse protein families with low sequence identity and shows superior performance in challenging biophysical contexts such as the transmembrane core and interfacial regions. Its validated computational efficiency enables large-scale mutation screening in minutes, offering a practical, robust, and powerful tool for transmembrane protein variant evaluation and engineering.
Allopolyploidy, involving whole genome duplication (WGD) of interspecific hybrids, is a driving force in the evolution of angiosperms, and has provided favored substrates for the domestication of major agricultural crops. This suggests allopolyploidy is a rich source of genetic variation amenable to natural and artificial selection. While allopolyploidy-induced chromosomal variation is common, its immediate phenotypic effects are challenging to delineate due to the confounding influence of postpolyploidy evolution. Newly constructed allopolyploids, having not yet undergone evolution, present suitable systems to address this issue. In this study, we synthesized five sets of allotetraploids, each with a unique genome constitution of S*S*DD, comprising a common paternal (DD) but distinct maternal (S*S*) parental diploid species of Aegilops. We observed that, except for one sterile synthetic allotetraploid, the remaining four allotetraploids exhibited high fertility, enabling the establishment of sexual lineages through selfing. Chromosomal variation in both number and structure occurred extensively, demonstrating moderate (though variable) effects on key morphological traits related to growth, development, and reproductive fitness of the nascent allotetraploids. All four sets of fertile allotetraploids can be crossed with bread wheat to generate pentaploid F1 hybrids, which as maternal parents can be further backcrossed to bread wheat. This approach promises a feasible strategy for the concomitant introgression of the vast repertoire of genetic variation from the D- and each of the four S* genome-containing species to bread wheat.
Precise modeling of RNA-ligand interactions is essential for understanding RNA functionality and designing RNA-targeted therapeutics. Current computational approaches largely focus on predicting discrete binding sites, limiting their applicability to complex RNA regions that may harbor multiple or diffuse ligand binding motifs. Here, we present RLAgent, an interactive agent framework designed to predict ligand interactions at the RNA region level, enabling higher-resolution and more flexible modeling than conventional site-centric approaches. RLAgent reframes the RNA-ligand prediction workflow as a dialogue-driven process. Through a natural language interface, users can interactively configure modeling preferences without writing code. A locally hosted large language model (LLM) acts as the core orchestration agent, automating all key components of the modeling pipeline, including data validation, feature encoding, model training, evaluation, and visualization. This agent-based design lowers technical barriers and enhances reproducibility, making RNA-ligand prediction more accessible for both computational and experimental researchers. ### Competing Interest Statement The authors have declared no competing interest.
Protein-protein interaction (PPI) variations are widely observed in many principal biological processes related to various diseases, while genetic mutation is one of the most common factors resulting in those variations. Unraveling these underlying biomolecular interactions is a key challenge that hinders a profound understanding of disease mechanisms. Although current models can accurately predict quantitative binding affinity, their application in pathogenic research is limited by difficulties in defining trend types, particularly extreme “disrupting” cases, along with poor generalizability beyond single-point mutations (SNP) and insufficient biological interpretability. Here, we present TAPPI, an interpretable large language model-driven deep learning framework, to perform Trends Assessment of Protein-Protein Interaction. It categorizes PPI variations’ trends into four types: disrupting, decreasing, no effect, and increasing, enabling fine-grained functional assessments and bridging genetic mutation and disease mechanisms. The framework generalizes effectively across both single-point and multi-point mutations and captures complex trend outcomes relevant to diverse disease landscapes. In benchmark evaluations, TAPPI achieved the state-of-the-art prediction performance and demonstrated biological relevance through interpretable results. In external datasets from Autism Spectrum Disorder, Lennox-Gastaut syndrome, and Epileptic Encephalopathies, given the premise that TAPPI’s predicted results align with pathogenic patterns, these results reveal how genetic mutations indirectly perturb disease-associated pathways and supporting our perspective of “crucial disrupting” and “bridging mutations-disease”. In summary, accurate, interpretable, and pathogenically relevant prediction of PPI variations trend predicted by TAPPI provides a novel approach for mechanistic discovery. ### Competing Interest Statement The authors have declared no competing interest. This work was supported by the National Natural Science Foundation of China (No: 62372099); the Jilin Scientific and Technological Development Program (No. 20230401092YY); Key Laboratory of Intelligent Rehabilitation and Barrier-free for the Disabled (Changchun University), Ministry of Education (2024KFJJ003)., No: 62372099, No. 20230401092YY, 2024KFJJ003