Prediction of pharmacokinetic (PK) properties is essential for early drug candidate screening and dosage regimen optimization. In recent years, using machine learning/deep learning approaches for predicting pharmacokinetic properties directly from chemical structures has attracted increasing attention. In this study, we propose multifidelity pharmacokinetic learning (MFPK), a transfer-learning framework for predicting intravenous pharmacokinetic parameters across multiple species, including humans, dogs, monkeys, rats, and mice. MFPK incorporates graph-, motif-, and three-dimensional structure-based molecular representations to capture comprehensive, multiscale chemical information. Comparative evaluations demonstrate that MFPK outperforms baseline models across multiple tasks, particularly volume of distribution at steady state (VDss) across all species (root-mean-square of logarithmic error (RMSLE) < 0.48, geometric mean fold error (GMFE) < 2.3). Furthermore, interpretability analyses were conducted to provide insights into model decision-making and mitigate the black-box nature of deep learning models. The MFPK model is accessible at https://lmmd.ecust.edu.cn/MFPK.
Complex diseases are commonly characterized by dysregulation across multiple targets and pathways, making drug combinations an important therapeutic strategy. This creates a growing need for computational methods to prioritize candidate combinations, yet current methods face several practical challenges, including cross-disease application, unified prediction of dual- and multidrug combinations, lack of reliable negative labels, and prediction for low-resource diseases. In this study, we proposed HyperDC, a unified cross-disease drug combination recommendation framework. HyperDC constructs a nonuniform hypergraph based on clinical and knowledge-driven drug-disease association data, representing single drugs and dual- and multidrug combinations in a unified modeling space. It further integrates knowledge graph pretraining and adversarial negative sampling to enhance model discrimination across tasks with different difficulty levels. In the unified dual-drug benchmark and disease-specific prediction tasks, HyperDC outperformed representative methods by up to 12.8 and 30.0 percentage points, respectively. In fixed-anchor clinical ranking tasks, HyperDC showed superior performance across multiple top-ranked recall settings and preferentially recalled FDA-approved dual-drug combinations and clinically supported three-drug combinations. In the data-sparse metabolic dysfunction-associated steatohepatitis (MASH) scenario, HyperDC completed a full workflow from anchor-drug identification and combination partner prioritization to in vitro experimental validation: 70% of the top 10 single-drug candidates were supported by recent literature, and 4 of 6 tested combinations showed synergistic lipid-lowering and anti-inflammatory effects, yielding a hit rate of 66.7%. Overall, HyperDC may narrow the screening space, improve prioritization efficiency, and provide methodological support for the development of combination therapy strategies for complex diseases.
Obesity is a major risk factor for type 2 diabetes. Some traditional Chinese medicine (TCM) compounds can improve obesity by increasing energy consumption. This study aimed to identify and verify TCM compounds that promote adipose thermogenesis to alleviate obesity via transcriptomic analysis, molecular docking, network pharmacology, and Connectivity Map (CMap) analysis. Thirty-six TCMs related to adipose tissue thermogenesis were first collected to generate transcriptomic data. Then a ranking method based on integrated transcriptomic data was used to select emodin from the TCM Da Huang for network pharmacology to explore its potential mechanisms and further experimental validation. CMap analysis of thermogenesis signature genes from ProFAT and GEO datasets identified triptolide as another compound for improving obesity through enhanced thermogenesis. In vitro experiments with 3T3-L1 cells and in vivo experiments with zebrafish were conducted. The results showed that both emodin and triptolide could upregulate the NAD+/NADH ratio, increase AMPK and SIRT1 expression, suppress lipid accumulation, and promote lipolysis, as verified by both in vitro and in vivo experiments. They also upregulated thermogenesis-related genes (e.g., UCP1, PGC1α, and PRDM16), lipolysis-related genes (e.g., PKA, ATGL, and HSL), mitochondrial biogenesis-related genes (e.g., NRF1, NRF2, and TFAM), and β-oxidation-related genes (e.g., CPT1α and CPT1β). Meanwhile, they downregulated lipogenesis-related genes such as FASN, SREBP1, and ACC. These results provide a basis for understanding the effects of emodin and triptolide on obesity. The findings highlight the effectiveness of combined screening approaches in exploring compounds with beneficial metabolic effects.
The catalytic activity of enzymes is highly dependent on the environmental pH. Mining enzymes with high activity under specific pH conditions can enhance catalytic efficiency, rendering significant industrial application value and economic benefits. To address the current challenge of insufficient accuracy in predicting enzyme optimal pH (pHopt), we developed active site‐based pHopt (AS‐pHopt), a prediction model enhanced by information of active site and pseudo‐label prediction. AS‐pHopt integrates key structural information from active sites and other physicochemical properties, which significantly influence enzyme pHopt. It uses Evolutionary Scale Modeling (ESM)‐2 with active site weighting to capture the active site‐specific information and ESM‐Cambrian to capture the overall protein representation. Additionally, a double cross‐attention mechanism is employed to merge both global and local protein information. To reduce prediction bias, we introduce a secondary loss based on cosine similarity and an label distribution smoothing‐weighted loss to optimize the model's overall performance. Compared with other existing studies, AS‐pHopt improves the R 2 by 8%–20%. Meanwhile, by combining few‐shot fine‐tuning, AS‐pHopt has achieved the screening and prediction of enzyme mutants to a certain extent. Overall, AS‐pHopt offers a promising strategy for accurate pHopt prediction and for advancing enzyme engineering across diverse biocatalytic applications.
Accurate prediction of a compound's site(s) of metabolism (SoMs) mediated by cytochromes P450 (CYP450) is advantageous in the early stage of drug discovery. However, existing computational methods often struggle to explicitly capture the microscopic electronic evolution associated with bond cleavage, and conventional graph neural networks face inherent challenges of information attenuation and oversmoothing during message passing, which restrict their ability to model long-range spatial dependencies, thereby limiting their prediction accuracy and generalization ability. To address these limitations, we propose a novel deep learning framework, CypGEM, based on a geometry-aware and edge-enhanced graph transformer for SoM prediction. By introducing the gated edge fusion and dynamic edge update mechanisms, CypGEM captures the microscopic electronic evolution features associated with chemical bond variations during metabolic reactions. Meanwhile, by integrating a global geometry-aware layer containing graph-wide topology and three-dimensional (3D) spatial information, CypGEM reconstructs long-range intramolecular steric constraints. Built on a constructed high-quality benchmark data set, CypGEM achieves better performance compared with existing models. Notably, in the case studies concerning FDA-approved drugs from recent years, the model demonstrates robust predictive performance when confronted with novel scaffolds unseen in the training set, exhibiting strong generalization ability. Furthermore, interpretability analysis confirms that the model has captured the synergistic rules of electronic effects and steric hindrance, providing medicinal chemists with structural optimization guidance grounded in physicochemical intuition. CypGEM is freely available at https://lmmd.ecust.edu.cn/CypGEM.
Antimicrobial resistance poses a significant challenge to conventional antibiotics, underscoring the urgent need for alternative therapeutic strategies. Antimicrobial peptides (AMPs) have emerged as promising candidates due to their broad-spectrum antibacterial activity and distinct mechanisms of action. This study presents ANIA, a deep learning framework developed to predict the minimum inhibitory concentration (MIC) values of AMPs against three clinically significant bacteria: Staphylococcus aureus, Escherichia coli, and Pseudomonas aeruginosa. ANIA leverages Chaos Game Representation (CGR) to transform AMP sequences into frequency-based image features, which are subsequently processed through a hybrid architecture comprising stacked Inception modules, a Transformer encoder, and a regression head. This integrative architecture enables ANIA to capture both local motif-based features and global contextual patterns embedded within AMP sequences. In benchmarking experiments, ANIA achieved notably superior performance compared to existing tools, including ESKAPEE-Pred, AMPActiPred, and esAMPMIC, achieving higher correlation coefficients and lower predictive errors across all bacteria targets, with the most pronounced improvement observed for P. aeruginosa, a pathogen renowned for its multidrug resistance. Specifically, ANIA achieved PCCs of 0.75-0.79 and MSEs of 0.23-0.26 across all species. Furthermore, motif-based interpretability analyses combining Grad-CAM visualizations, correlation heatmaps, motif frequency distributions, and hydrophobicity profiling revealed biologically meaningful subregions within the CGR matrix that are plausibly associated with antimicrobial efficacy. In conclusion, this study develops ANIA as a robust predictive tool for MIC estimation, offering valuable insights into the design of effective antimicrobial agents and contributing to the fight against antimicrobial resistance. A user-friendly web server for ANIA is available at https://biomics.lab.nycu.edu.tw/ANIA/.
Clinical attrition in drug development is frequently driven by suboptimal pharmacokinetic and toxicological (ADMET) properties rather than inadequate target efficacy. Accordingly, early-stage ADMET assessment has become an increasingly important component of Structure-Based Drug Design (SBDD). Here, we present empirically derived chemical navigability guidelines based on an analysis of ADMET-related properties predicted by admetSAR 3.0 across approved drugs curated from DrugBank. This study extracts interpretable patterns from model-derived data to support early-stage compound prioritization. Threshold ranges were established for 117 endpoints, including drug-induced liver injury, hERG inhibition, mutagenicity, intestinal absorption, cytochrome P450 interactions, and blood–brain barrier permeability. These parameters were integrated into an intuitive color-coded visualization framework for rapid compound assessment. The proposed guidelines are intended as context-dependent heuristics derived from statistical trends within the predicted chemical space of approved drugs rather than as universal decision rules. The utility of the ADMET-first strategy was further evaluated using an external and independent library of 1,756 KEAP1/NRF2 modulators compounds from ChEMBL, employing experimentally determined biological activity (pChEMBL) instead of docking-derived metrics. Enrichment analysis demonstrated that ranking compounds according to the ADMET-score identified true active compounds substantially earlier than random selection, achieving an enrichment factor (EF) of 1.29 in the top 5% of the library and recovering approximately 80% of active compounds within a limited fraction of the evaluated chemical space. These findings support chemical navigability as an ADMET-driven framework for efficient early-stage compound prioritization and virtual screening. In addition, we provide a freely accessible, curated database of more than two million ADMET-annotated commercially available compounds from the MolPort library, prefiltered according to the proposed chemical navigability guidelines. All datasets and scripts are publicly available through admetSAR.umh.es.
The accurate identification of cytochrome P450 (CYP) substrates is crucial in drug discovery and safety assessment, as these enzymes mediate the metabolism of most clinical drugs. However, existing computational models are often limited by data quality issues and lack the ability to quantify prediction uncertainty, hindering their reliable application. To address these challenges, we present EviCYP, a novel prediction framework that integrates evidential deep learning with vector quantization (VQ). We first constructed a high-quality data set by curating 4388 substrates and 2880 nonsubstrates from 1629 publications, and supplemented it with 3728 pseudonegative samples, resulting in 10,996 samples spanning nine major CYP isoforms. The EviCYP architecture processes multimodal molecular representations and enzyme sequences through dedicated encoders, compresses features via VQ to reduce redundancy, and employs an evidential layer to output both class probabilities and an uncertainty estimate. On an internal test set, EviCYP achieved an average AUROC of 0.9500. Notably, the model's uncertainty quantification is highly reliable, with high-uncertainty predictions strongly correlating with classification errors. This work provides a robust and trustworthy computational tool for CYP substrate prediction.
Recent studies have revealed that some noncoding RNAs (ncRNAs) bear translational potential, and their encoded micropeptides have essential functions in multiple biological processes. However, accurate identification of coding-capable ncRNAs remains challenging due to weak translation signals, low conservation, and heterogeneous data distributions. Herein, we propose ncProFormer, a deep learning framework tailored for ncRNA coding-potential prediction. ncProFormer integrates the nucleic-acid language model GENA-LM to obtain contextual sequence embeddings, adopts an all-token representation strategy, and employs a convolutional neural network (CNN)-enhanced transformer encoder to jointly capture local nucleotide patterns and long-range dependencies. ncProFormer consistently outperformed the existing methods across the in-house human data set, the external validation data set, the public CPPred benchmark data set. More importantly, this study presents the first cross-species evaluation in ncRNA coding-potential prediction. Without retraining, ncProFormer maintained its strong predictive performance on mouse and rat data sets, showing that the learned biological representations are transferable and it is robust under the distributional shift and cross-species conditions. Collectively, these findings establish ncProFormer as an effective and generalizable framework for uncovering the coding potential of ncRNAs, thus offering a promising computational tool for characterizing ncRNA functions across diverse transcriptomic contexts.
BACKGROUND:Women face a heightened risk of Alzheimer's disease (AD), partly attributed to post-menopausal estrogen loss. Given that ERβ activation avoids the oncogenic risks of ERα and GPR40 plays a pivotal role in neuronal function, the ERβ/GPR40 axis show a promising therapeutic target for anti-AD drug discovery. To inspect the role of this axis, we employed Vincamine (Vin), a monoterpenoid indole alkaloid from Madagascar periwinkle that we previously identified as a GPR40 agonist. PURPOSE:To elucidate the role of ERβ/GPR40 axis in AD pathogenesis and to investigate the therapeutic potential of Vin in ameliorating AD-related deficits. METHODS:We combined analyses of clinical data from female AD patients (GSE33000) with the research in 3×Tg-AD mice to examine the differences in ERβ/GPR40 expression. The binding of ERβ and GPR40 was detected by CUT&Tag assay, protein-DNA docking simulation and molecular dynamics simulation assays. Vin was used to evaluate the therapeutic potential of ERβ/GPR40 axis activation for AD. The underlying mechanisms were investigated by assay against the adeno-associated virus (AAV)-CMV-PHP.eB-KD-GPR40 injected 3×Tg-AD female mice. RESULTS:ERβ and GPR40 are both downregulated in brains of female AD patients and 3×Tg-AD mice, and ERβ directly binds to GPR40 promoter. Brain-specific GPR40 knockdown caused cognitive impairment in female wild type (WT) mice. Vin as a GPR40 agonist but not an ERβ ligand ameliorated AD-like pathology in 3×Tg-AD female mice. Specifically, Vin suppressed neuroinflammation via GPR40/NF-κB/NLRP3 pathway, inhibited neuronal tau hyperphosphorylation via GPR40/GSK3β/CaMKII pathway, while promoted synaptic plasticity via GPR40/PKA/CREB/BDNF pathway. CONCLUSION:To our knowledge, our study provides the first identification of the specific ERβ-binding regions and key residues within the GPR40 promoter, offering novel mechanistic insight into their transcriptional regulation. Furthermore, our work establishes ERβ/GPR40 axis as a potentially therapeutic strategy for female AD and highlight the medication interest of Vin in treating this disease.
Prediction of cytochrome P450 (CYP) induction is highly advantageous in early stage drug discovery, as it helps mitigate the risks of drug-drug interactions and toxicity. However, the development of specialized predictive models for CYP induction remains limited, largely due to the scarcity of available inducer data. To address these challenges, we propose ULCYP, a multitask deep learning framework based on positive-unlabeled (PU) learning for the prediction of CYP induction. ULCYP effectively leverages large-scale unlabeled data to compensate for the lack of trustworthy negative samples, thereby enabling a more accurate estimation of the decision boundary. Comparative evaluations demonstrate that ULCYP outperforms baseline models across multiple performance metrics, achieving an average AUC greater than 0.81 on the test set comprising CYP inducers and nonagonists of key CYP induction mediators, including the pregnane X receptor (PXR), constitutive androstane receptor (CAR), and aryl hydrocarbon receptor (AhR). To enhance prediction reliability and interpretability, the integrated gradients method was employed to elucidate key molecular substructures driving model predictions, complemented by a rigorously defined applicability domain. The ULCYP model is publicly accessible at https://lmmd.ecust.edu.cn/ULCYP/.
Clinical attrition in drug development is frequently driven by suboptimal pharmacokinetic and toxicological (ADMET) properties rather than inadequate target efficacy. Accordingly, early-stage ADMET assessment has become an increasingly important component of Structure-Based Drug Design (SBDD). Here, we present empirically derived chemical navigability guidelines based on an analysis of ADMET-related properties predicted by admetSAR 3.0 across approved drugs curated from DrugBank. These parameters were integrated into an intuitive color-coded visualization framework for rapid compound assessment. The proposed guidelines are intended as context-dependent heuristics derived from statistical trends within the predicted chemical space of approved drugs rather than as universal decision rules. The utility of the ADMET-first strategy was further evaluated using an external and independent library of 1,756 KEAP1/NRF2 modulator compounds from ChEMBL, employing experimentally determined biological activity (pChEMBL) instead of docking-derived metrics. Enrichment analysis demonstrated that ranking compounds according to the ADMET-score identified true active compounds substantially earlier than random selection, achieving an enrichment factor (EF) of 1.29 in the top 5% of the library and recovering approximately 80% of active compounds within a limited fraction of the evaluated chemical space. We provide a freely accessible (https//admetSAR.umh.es), curated database of more than two million ADMET-annotated commercially available compounds from the MolPort library, prefiltered according to the proposed chemical navigability guidelines.
Near-infrared (NIR) genetically encoded fluorescent tags are superior for new imaging capabilities ranging from multiplexed imaging to deep-tissue imaging. Here, we describe the development of NirFAP680; NirFAP680 is an NIR fluorogen-activating protein (FAP) that consists of a newly designed NIR fluorogen termed HBMT and an engineered protein tag termed NirFAP. NirFAP680 has an order of magnitude greater cellular brightness and superior photostability compared to currently available NIR FAPs and fluorescent proteins in both single- and two-photon excitation and allows robust imaging of proteins in live cells and in vivo. Owing to its unique spectroscopic property of a >100 nm Stokes shift, NirFAP680 can also serve as a superb receptor for fluorescence or bioluminescence resonance energy transfer, allowing real-time monitoring of protein-protein interactions in live cells and sensitive bioluminescence imaging in vivo, respectively. Therefore, NirFAP680 will likely be useful for advanced biological imaging in live cells and in vivo.
UDP-glucuronosyltransferases (UGTs) play a critical role in drug metabolism by catalyzing the glucuronidation of structurally diverse compounds. However, accurately predicting UGT-mediated sites of metabolism (SOMs) remains a challenge due to the limited availability of annotated data. In this study, we introduce UGTformer, a unified graph transformer-based framework that simultaneously performs UGT substrate classification and SOM prediction. UGTformer employs a hierarchical architecture integrating multi-hop message propagation with hop-aware and node-level transformer encoders. The model was pretrained on large-scale molecular graphs via chemically informed self-supervised tasks, and fine-tuned on a manually curated UGT metabolism data set covering four major metabolic reaction categories. In five-fold cross-validation, UGTformer achieved an AUC of 0.833 for substrate classification and 0.884 for SOM identification, outperforming multiple GNN baselines. On an independent external validation set, it maintained robust performance, demonstrating strong generalization to previously unseen molecules. By integrating chemically meaningful structural encodings and a joint learning paradigm, UGTformer delivers interpretable and biologically consistent predictions, offering a reliable and scalable approach for UGT-related metabolism prediction. The UGTformer model is freely accessible at https://lmmd.ecust.edu.cn/UGTformer/.
Prediction of intrinsic clearance (CLint), a key parameter of metabolic stability, is critical for pharmacokinetic assessment and early drug candidate screening. However, existing predictive models for CLint are often constrained by the quality of publicly available data and insufficient biological interpretability. In this study, CLint data were systematically collected from public databases and manually verified against original experimental records. A pragmatic filtering step was employed as part of the data curation process to mitigate the influence of potentially inconsistent measurements. Based on the curated data set, we developed traditional machine learning (ML) models and a dual-branch deep learning (DL) architecture that integrates molecular fingerprints with graph-based structural features. Furthermore, we proposed an ensemble strategy that dynamically combines ML and DL predictions according to molecular similarity. The ensemble model achieved the best overall performance, with an R2 of 0.634 on the test set. To elucidate the biochemical determinants underlying these predictions, we conducted interpretability analyses that linked model outputs to molecular physicochemical properties and potential metabolic sites, revealing a complementary representation pattern between the ML and DL models. Together, the proposed modeling framework and its mechanistic insights provide a biologically informed tool for early pharmacokinetic screening and contribute to a deeper understanding of structure-metabolism relationships. An interactive web server has been developed to facilitate the practical application of the proposed model and is publicly available at https://lmmd.ecust.edu.cn/clint/.
Abstract The global antibiotic resistome remains largely unexplored, not because antibiotic resistance genes (ARGs) are rare in the environment, but because many are evolutionarily distant from known ARGs. Current computational approaches primarily rely on sequence homology, and thus miss distant homologues. We develop GeoARG, a geometry-enhanced framework that integrates structural features with protein language models through knowledge distillation, enabling efficient large-scale screening using sequence input alone. Across multiple benchmarks, GeoARG substantially improves the detection of remotely homologous ARGs, particularly under low sequence identity and fragmented conditions. Large-scale metagenomic analysis uncovers 1,485 high-confidence ARG candidates that are highly divergent from known ARGs, expanding the phylogenetic and functional landscape of the resistome. Structural analyses further show that these candidates preserve active-site geometry and maintain stable ligand-binding configurations consistent with known resistance mechanisms. These results demonstrate that geometric constraints enable systematic expansion of the resistome and facilitate the discovery of evolutionarily distant yet functionally conserved genes. A public web server is available at https://ycclab.cuhk.edu.cn/GeoARG/ .
[This corrects the article DOI: 10.1016/j.jpha.2025.101317.].
Lung adenocarcinoma (LUAD), the most common subtype of nonsmall cell lung cancer, exhibits substantial molecular heterogeneity, complicating subtype classification, progression assessment, and treatment decision-making. Advances in high-throughput sequencing enable multi-omics analysis to reveal cancer mechanisms and biomarkers, yet the high dimensionality, heterogeneity, and interrelationships of omics layers such as transcriptome, microRNA expression, methylome, and copy number variation remain challenging to integrate through conventional methods. Most existing graph-based approaches represent patients as nodes, obscuring gene-level regulatory dynamics and limiting biological interpretability. To address this, we propose the Multi-omics Hierarchical Graph Neural Network (MoAGNN), a novel architecture that represents genes as nodes, integrates four omics, and leverages graph convolution with self-attention-based graph pooling to identify informative molecular nodes, thereby enhancing predictive performance and interpretability for LUAD subtype classification, tumor staging, and prognosis prediction. Multi-omics datasets from The Cancer Genome Atlas (TCGA) were used and results showed that MoAGNN achieved a test accuracy of 0.89 for LUAD subtype classification, outperforming conventional models (Random Forest, Support Vector Machine and Multi-Layer Perceptron) as well as state-of-the-art graph-based models MoGCN, a multi-omics integration model based on graph convolutional network, and MOGLAM, an end-to-end interpretable multi-omics integration method. Furthermore, we validated the generalizability of this framework on the GSE81089 dataset, demonstrating its potential applicability to clinically relevant risk assessment. Subsequent functional enrichment and survival analyses validated the biological relevance of the key genes identified by MoAGNN, supporting their potential roles in LUAD progression, and suggesting the broader applicability of this framework in multi-omics cancer research.