
The emergence of multidrug-resistant (MDR) and extensively drug-resistant (XDR) Mycobacterium tuberculosis strains highlights the requirement for new therapeutic targets and antitubercular chemotypes. CysA2, a sulfur transferase that participates in sulfur metabolism and redox homeostasis, is a relatively underexplored therapeutic target. Here we used an integrated computational approach to discover potential small molecule inhibitors of CysA2. This led to prioritization of 448 candidates for MM-GBSA rescoring from the > 448,000 compounds screened, where the binding free energies ranged from -68.57 to -14.10 kcal/mol, compared to -16.25 kcal/mol for the reference ligand S-nitrosoglutathione (GSNO). The best candidates were F0808-0813, F1631-0176 and F2532-0240 with MM-GBSA binding energies of -68.57, -62.98 and -61.46 kcal/mol, respectively. Further 500 ns molecular dynamics simulations showed particularly favorable dynamic behavior for F0808-0813 and F2532-0240, while F1631-0176 displayed comparatively higher ligand mobility and conformational variability. F2532-0240 displayed the most favorable overall dynamic profile, including stable conformational sampling and the lowest superimposition RMSD (0.771 Å) of the investigated systems. The DFT analysis further confirmed the electronic stability of F0808-0813 and F2532-0240, with HOMO-LUMO gaps of 4.517 and 4.000 eV, respectively, versus 3.025 eV for GSNO. QM/MM analysis further supported accommodation of prioritized ligands in the binding environment of CysA2. Taken together, the results point to F2532-0240 and F0808-0813 as the top CysA2-directed candidates for further experimental validation and structure-guided anti-tubercular drug development.
Motivation:RNA molecules play critical roles in gene regulation, viral replication, and cellular control, with their functions tightly coupled to three-dimensional structure. Advances in cryogenic electron microscopy (cryo-EM) now enable RNA structure characterization across a broad resolution range. RNA secondary structural motifs, including hairpins, internal loops, and bulges, act as fundamental building blocks of RNA tertiary architecture and are key targets in RNA-focused therapeutic design. Despite this, most computational approaches for RNA structure prediction from cryo-EM density maps do not explicitly utilize secondary structural motifs as intermediate representations, largely due to the absence of large-scale, high-quality, and motif-resolved datasets suitable for machine learning. Results:Here, we present a large, open-source dataset containing over 125,000 motif-resolved cryo-EM density maps paired with corresponding atomic structures, spanning 25 classes of RNA secondary structural motifs. The dataset covers resolutions from 1.5 Å to 34.0 Å, encompassing both near-atomic and low-resolution density maps relevant to RNA modeling. Each motif instance includes a segmented cryo-EM density map represented as a standardized 3D voxel grid, with atomic-level motif annotations propagated to voxel-level labels for RNA backbone, ribose sugar, and nucleobase components. Segmentation quality is validated via cross-correlation analysis, demonstrating strong agreement between motif-level density maps and atomic reference models. To illustrate the dataset's utility, high-resolution maps (1.5-2.8 Å) were used to train a machine learning classifier that distinguished five motif classes with a specificity of 0.948. Availability and Implementation:Source code, implementation of the fully automated pipeline, and the benchmark datasets are publicly available at. GitHub:https://github.com/DrDongSi/3DEM-RNA-Motif-Dataset. Zenodo:https://zenodo.org/communities/3dem-rna-motif-dataset.
MOTIVATION:RNA functions in gene regulation, viral replication, and cellular control are tightly coupled to three-dimensional structure and local conformational features. Cryogenic electron microscopy (cryo-EM) now enables RNA structure characterization across a broad resolution range, but full maps are large, heterogeneous, and variable in local resolution. RNA secondary structural motifs, including hairpins, internal loops, and bulges, provide recurring local units for interpreting RNA density, comparing structures, and developing machine-learning models. Existing cryo-EM-based methods generally focus on complete maps, chains, residues, or atomic model construction rather than motif-level representations, partly because large-scale motif-resolved cryo-EM datasets remain limited. RESULTS:We present an open-source dataset of more than 100,000 motif-resolved cryo-EM density segments paired with atomic structures, spanning 25 RNA secondary structural motif classes and resolutions from 1.5 Å to 34.0 Å. Each motif is represented as a standardized 3D voxel grid with voxel-level labels for RNA backbone, ribose sugar, and nucleobase components. Motif-level map-model agreement was evaluated using masked cross-correlation (CCmask) and atom-level Q-scores, revealing resolution-dependent trends in regional density agreement and atomic resolvability. As a baseline benchmark, a 3D convolutional neural network trained on a curated, class-balanced, primarily high-resolution subset distinguished five motif/background classes, achieving macro-averaged sensitivity of 0.836 ± 0.019, specificity of 0.958 ± 0.005, balanced accuracy of 0.897 ± 0.012, and G-mean of 0.894 ± 0.013. AVAILABILITY AND IMPLEMENTATION:Source code, pipeline implementation, benchmark datasets, and an interactive web application are available at GitHub (https://github.com/DrDongSi/3DEM-RNA-Motif-Dataset), Zenodo (https://zenodo.org/communities/3dem-rna-motif-dataset), and Hugging Face Spaces (https://huggingface.co/spaces/houlab/arsma-cryoem).
Matrix metalloproteinase-12 (MMP-12) is a zinc-dependent endopeptidase that plays an important role in the pathogenesis of several inflammatory, pulmonary, cardiovascular, neurological, and cancer-associated disorders. Despite its therapeutic significance, the development of potent and selective MMP-12 inhibitors remains challenging because of the high structural similarity shared among various MMP family members. This study presents an integrated framework combining matched molecular pair (MMP) cliff analysis, quantitative read-across structure-activity relationship (qRASAR)- based modelling, deep learning-based analysis of binding interactions, and MD simulations to elucidate the structural determinants governing potent MMP-12 inhibition, with particular emphasis on recognition of the S1' pocket. Leveraging structural similarity with MMP-12 inhibitors, the final qRASAR MLR model showed satisfactory predictive performance (R2 = 0.710, Q2F1 = 0.734, Q2F2 = 0.734, and MAEtest = 0.563). A physics-aware, deep learning-based binding interaction analysis showed that potent inhibitors like C5 and C54 form favourable interactions with key residues in the S1' pocket, including P238, Y240, K241, and F248, whereas weak inhibitors like C475 exhibit comparatively weaker engagement within this S1' subsite. Subsequently, MD simulations further confirmed the enhanced stability, compactness, and reduced conformational flexibility of the MMP-12-C5 and MMP-12-C54 complexes relative to MMP-12-C475 and highlighted the critical role of persistent S1' pocket interactions in stabilizing the protein-ligand complexes and enhancing MMP-12 inhibitory potency. The findings emphasize the importance of effective S1' pocket recognition and provide valuable insights for the rational design of potent MMP-12 inhibitors in future.
The prediction of conformational B-cell epitopes (BCEs) is crucial for vaccine development and therapeutic antibody design. However, reliable identification of BCEs remains challenging because epitope residues are spatially discontinuous and represent only a small fraction of antigen surface residues, leading to severe class imbalance and high false-positive rates. In this study, we propose GCAT-BCE, a hybrid graph neural network that integrates graph convolutional networks (GCN) and graph attention network (GAT) for conformational BCE prediction. To reduce prediction noise, buried residues are first removed through a relative solvent accessibility (RSA)-guided filtering strategy prior to graph construction. Then, GCAT-BCE leverages multi-modal residue-level features (including amino acid types, secondary structure, relative solvent accessibility, and epitope propensity) combined with three stacked GCN layers with residual connections to capture local spatial interactions, followed by a GAT layer to refine long-range residue dependencies. Comprehensive evaluations on two independent benchmark test sets comprising 15 and 45 antigens demonstrated that GCAT-BCE consistently outperformed state-of-the-art sequence-based and structure-based models. Notably, GCAT-BCE achieved the highest AUC-PR and AUCPR10% values, indicating superior capability in identifying true epitope residues among highly imbalanced samples. Furthermore, we evaluated the GCAT-BCE model on two newly curated independent test sets with 215 non-redundant antigens. The results demonstrated that GCAT-BCE consistently outperformed BepiPred-3.0, CALIBER, BIDpred, and CLBTope with particularly pronounced improvements in AUC-PR. On the RoBep_187 test set, GCAT-BCE achieved an AUC-PR of 0.392, approximately twice that of the second-ranked model, while on the PDB2526_28 test set it maintained the highest AUC-PR and AUCPR10% performance, highlighting its robust predictive capability for minority-class epitope residues.
Computational prediction of PROTAC degradation activity (DC50) has attracted growing interest, yet the reliability of reported model performance remains poorly understood because sufficiently stringent evaluation protocols are rarely applied. Here, we present a hierarchical benchmark designed to expose evaluation pitfalls and quantify the transferability and reliability limits of current PROTAC predictors. Using a curated dataset of 2405 DC50 measurements spanning 22 target proteins and two E3 ligases (CRBN and VHL), we benchmarked classical machine learning (Random Forest, ExtraTrees, Ridge, PLS), gradient-boosted trees (XGBoost), nearest-neighbor retrieval baselines, protein negative controls, and a representative multi-modal deep learning ensemble (HybridMoECrossAttn) across Random, Scaffold, Leave-One-Target-Out (LOTO), and Leave-One-Family-Out (LOFO) splits. Under Random evaluation, a simple Random Forest + ECFP4 baseline achieved pooled R² = 0.693 ± 0.025, indicating that conventional models already approach the apparent ceiling under interpolation-oriented settings. However, all methods collapsed under LOTO (best R² = -0.012), revealing that much of the apparent progress in the literature reflects chemical-neighbor memorization rather than robust target-level generalization. We further show that target-wise error is significantly associated with continuous protein semantic proximity in ProtBERT space (Spearman ρ = -0.461, p = 0.047), whereas coarse family-level descriptors are uninformative. A four-quadrant failure taxonomy reveals that protein shift is more damaging than chemical novelty (MAE 1.09-1.12 vs. 0.85-0.97), and conformal prediction becomes severely overconfident under target extrapolation, with empirical 90% coverage dropping to 63.9-66.7%. These results reposition PROTAC prediction as a problem of transferability and reliability rather than leaderboard optimization and provide practical guidelines for future benchmark design.
The Hedgehog (Hh) signaling pathway, crucial for embryonic development and tissue homeostasis, is frequently dysregulated in cancers, making its core component, the Smoothened (SMO) receptor, a prime therapeutic target. This study employed an integrative computational strategy to identify novel SMO inhibitors. Virtual screening of the ZINC database highlighted ZINC000299767263 (ZINC7263) as a top candidate, exhibiting a Glide score of -10.86 and MM/GBSA binding energy of -65.3 kJ/mol, comparable to the FDA-approved inhibitor Vismodegib. Molecular docking revealed critical interactions with SMO residues Asp384, Arg400, and Gln477, supported by hydrophobic contacts with Phe484 and Trp480. Molecular dynamics simulations (200 ns) demonstrated stable binding, with root-mean-square deviation (RMSD) and fluctuation (RMSF) analyses confirming complex stability. Free energy decomposition underscored contributions from Asp384, Arg400, and Gln477, aligning with Vismodegib's binding mechanism. ZINC7263 adhered to drug-likeness criteria (Lipinski's rule of five, Jorgensen's rule of three) and exhibited favorable pharmacokinetics, including high oral absorption (89.45%) and blood-brain barrier permeability (logBB = -1.01). Induced-fit docking further validated its binding mode, highlighting a conserved interaction profile. These findings position ZINC7263 as a novel, potent SMO inhibitor with a stable binding affinity, dynamic interaction stability, and promising preclinical potential for targeting Hh-driven cancers. Its unique scaffold and optimized pharmacokinetic profile advocate for advanced in vitro and in vivo studies to translate these computational insights into therapeutic applications.
BACKGROUND:Multi-epitope-based peptide (MEBP) vaccines have emerged as promising alternatives to conventional vaccines for combating Mycobacterium tuberculosis (M.tb) infection. Current immunoinformatics-based MEBP vaccine design strategies typically construct a single vaccine candidate by assembling the selected epitopes in an arbitrary linear order, interspersed with peptide linkers and adjuvants. This conventional approach implicitly assumes that epitope order has little or no influence on vaccine performance and consequently characterizes only one construct from the vast combinatorial space of possible epitope arrangements. As a result, the majority of structurally distinct MEBP variants remain unexplored, precluding a systematic evaluation of how epitope positional rearrangement influences structural conformation, physicochemical and energetic properties, and predicted immunogenicity. This limitation substantially constrains the rational identification and optimization of more effective MEBP vaccine candidates against M.tb. OBJECTIVES:The present study aims to systematically evaluate the influence of epitope order (positional rearrangement) on the structural stability, energetic profile, and predicted immunological performance of multi-target tuberculosis multi-epitope-based vaccine candidates (MEBPVC) using a rationally developed and comprehensive comparative computational workflow. METHODS:Five target proteins were selected from an initial set of fifteen antigenic proteins of the M.tb H37Rv proteome based on their high antigenicity and minimal sequence similarity to the human proteome. From each target protein, the two top-ranked Dual-Trigger Epitopes (DTEs) were identified, yielding a total of ten DTEs. These DTEs were assembled, together with an adjuvant and appropriate peptide linkers, generating a representative MEBPVC. Subsequently, a comprehensive library comprising all 10! (3628,800) unique positional rearrangements from the above ten DTEs' set was computationally generated. The resulting MEBPVC variants were systematically screened and evaluated based on their predicted structural, physicochemical, energetic, and immunological properties to identify the most promising vaccine candidates. The top-ranked candidates were further characterized through molecular docking with TLR4, followed by 500 ns molecular dynamics simulations, MM/GBSA binding free energy calculations, immune response simulations, and in silico cloning to assess receptor-binding affinity, complex stability, immunogenic potential, and expression feasibility. RESULTS:Comprehensive evaluation of all 10! (3628,800) MEBPVC variants demonstrated that epitope order influences the structural, energetic, receptor-binding, and immunological properties of MEBPVCs. Systematic screening identified four top-ranked candidates, with M.Tb_MEBPVC_2595872 emerging as the lead construct. It achieved the highest influence score (3.15) and formed a structurally stable TLR4 complex characterized by low RMSD (0.58 nm), low RMSF (0.13 nm), persistent hydrogen-bond interactions (20-27), and favorable MM/GBSA binding free energy (-78.22 kcal/mol). Immune simulations predicted robust humoral and cellular immune responses, while in silico cloning confirmed compatibility with the pET expression system. CONCLUSION:The findings support our hypothesis that epitope order is a critical determinant of the structural, receptor-binding, and immunological properties of MEBP vaccine constructs and should therefore be considered an essential optimization parameter in computational vaccine design. Systematic generation and comparative evaluation of positional variants exposed the limitations of the conventional single-construct design strategy and enabled the identification of four top-ranked vaccine candidates, with M.Tb_MEBPVC_2595872 emerging as the most promising candidate against M.tb. Experimental validation is warranted to confirm the predicted immunogenicity and protective efficacy of these candidates.
An integrated computational strategy was used to identify potential inhibitors of the NAD (+)-arginine ADP-ribosyl transferase of cholera toxin. The target protein structure was obtained from the UniProt database and validated using PROCHECK. The potential binding pockets were predicted using DeepSite. A total of 50 flavonoids were screened using structure based virtual screening (VS), of which Kaempferol (-9.1 kcal/mol) and Taxifolin (-8.9 kcal/mol) were identified as potential inhibitors. Molecular interaction studies indicated that kaempferol had a substantially larger interaction network (greater number of hydrogen bonds, pi-pi interactions and hydrophobic interactions) than taxifolin. ADMET and toxicity profiling for each compound indicated that both compounds had favorable drug-like properties including high GI absorption, low BBB penetrability, and acceptable toxicity profiles. 100 ns molecular dynamics analyses showed stable binding for both ligands. Kaempferol had faster equilibration times, less flexible protein, and less solvent exposure than taxifolin. Additionally, MM-GBSA binding free energy analyses further corroborate the binding affinity of kaempferol/taxifolin was stronger (-36.28 vs -30.62 Kcal/mol binding affinity). Collectively, these findings show that kaempferol is a viable lead compound for cholera toxin suppression and serve as the foundation for future antitoxin medication development and experimental validation.
Pan-cancer multi-omics analysis requires models that integrate complementary molecular signals while preserving biologically meaningful relationships among patients. This study presents a leakage-controlled benchmarking framework for patient-graph learning in pan-cancer classification and prognosis analysis, focusing on how graph construction affects downstream performance. The benchmark explicitly separates fold-specific graph formation from downstream prediction. Using the TCGA Pan-Cancer cohort of 8204 primary tumors across 31 cancer types, RNA expression and copy-number variation data were used to compare early feature fusion, lightweight similarity network fusion (SNF-lite), and fused k-nearest neighbor similarity graphs under a common GATv2 encoder family with a matched attention-head search space and inner-validation selection procedure. A strict 5 × 3 nested cross-validation protocol ensured that imputation, gene selection, feature scaling, similarity computation, and neighbor search were fitted on training folds only. At G'=2000, graph-level fusion approaches achieved about 0.92 accuracy and 0.89 Macro-F1, outperforming early fusion at about 0.89 accuracy and 0.84 Macro-F1. Fused kNN graphs also showed higher neighborhood label purity than SNF-lite despite similar predictive performance. A weighted topology audit showed that local label agreement alone did not determine graph utility. Gene and omics ablations showed that RNA carried the dominant subtype-discriminative signal, while CNV and mutation contributed weaker but complementary information. A Cox auxiliary objective retained classification performance when used alone and enabled out-of-fold prognostic stratification. These findings show that patient-graph construction is a key design choice in pan-cancer multi-omics learning and that leakage-controlled evaluation is essential for reliable and biologically informative benchmarking in computational oncology.
Drug response prediction in acute myeloid leukemia (AML) is challenged by sample heterogeneity and high-dimensional RNA sequencing profiles. We develop NGM-AML, a modular model predicting ex vivo drug response from BeatAML2 RNA-seq and sensitivity data. After filtering, 306 Waves 1+2 samples (28055 records) serve as training and 173 Waves 3+4 samples (16123 records) as the test set across 111 drugs. Drug targets and AML prior genes are mapped to a PPI network; random walk with restart and community detection construct 45 modules. Per drug, module scores are partitioned into sensitivity and resistance components by their association with the area under the dose-response curve. The results show that NGM-AML achieves mean Pearson and Spearman correlations of 0.324 and 0.326 across drugs, with an MAE of 37.396. Pooling all test records yields a Pearson correlation of 0.703 between predicted and observed AUC. For representative drugs, Pearson correlations reach 0.780 for Venetoclax and 0.631 for Trametinib. Within patients, median Spearman correlation and NDCG@5 are 0.75 and 0.96 for drug ranking. Runtime decreases from 87125.4 s for the raw RNA-seq model to 1102.5 s for NGM-AML. Enriched processes include extracellular matrix adhesion, integrin signaling, and RTK/MAPK pathways, consistent with known AML survival and drug-resistance mechanisms.
BACKGROUND:Nasopharyngeal carcinoma (NPC) is a multifactorial disease driven by both genetic and environmental factors. The neurogenic locus notch homolog 1 (NOTCH1) gene has dual oncogenic and tumor-suppressive effects that depend on the cellular context. An intronic single nucleotide polymorphism (SNP) in NOTCH1, rs3124599, has been linked to various diseases. However, its involvement in NPC remains unknown. This study aims to explore the regulatory and susceptibility effects of rs3124599 on NPC. METHODS:Integrative computational analyses were performed using various bioinformatics tools and databases. These included RegSNPs-intron, PredictSNP, RegulomeDB, HaploReg v4.2, SNP2TFBS, atSNP, SNPnexus, eQTLGen, and GTEx. We used these resources to assess potential effects of rs3124599 on splicing, TF binding, chromatin accessibility, histone modifications, and gene expression. RESULTS:The intronic SNP rs3124599 is located in an epigenetically dynamic region. This region is enriched with active histone marks and DNase I hypersensitivity (DHS) sites. These features consistent with open chromatin. Motif analyses revealed disrupted binding affinities for significant TFs, including EGR2-4, PLAG1, NFE2, MYC, and FOXO3, indicating regulatory disruption. FunSeq2 predicted this SNP to be deleterious (score = 0.90). eQTL analysis validated its cis-regulatory function, revealing that it activated DNLZ, INPP5E, and SEC16A and repressed GPSM1 and CARD9. No significant findings were observed for lncRNA and miRNA. CONCLUSION:Despite its intronic location, SNP rs3124599 lies within an epigenetically regulated region that modifies NOTCH1 expression and NPC susceptibility at the chromatin/transcriptional level. Experimental validation is highly recommended to elucidate its role in NOTCH1 regulation and NPC progression.