
Abstract Retrosynthesis-based synthesizability scoring triages molecules from generative design but is expensive: every score requires a multi-step search. We present SynOmega, an open-source toolkit that couples a single-step template model, an AND–OR route search, and a route-based synthesizability score (SynScore). Its single-step model can be restricted, at the reaction-template level, to simplifying disconnections that split the target into smaller precursors. On 1000 ChEMBL drug molecules this yields two findings. (1) The simplifying constraint cuts node expansions by about 30% on jointly solved targets and lowers the median search time by about a third. This reduction in search effort is the robust, budget-independent result; the small accompanying rise in solved rate is a secondary effect between the two separately trained models, not the isolated result of toggling one model’s action space. (2) As a complete system, under matched search depth, width and iteration budget, SynOmega reaches about 1.8× the solved rate of the open-source planner AiZynthFinder while searching about 13× faster. SynOmega thus offers a cheap, data-level action-space constraint that makes route-based synthesizability scoring more efficient without sacrificing solvability.
Abstract Identification of RNA-small-molecule binding sites is a critical first step in RNA-targeted drug discovery. Although several machine learning methods have made progress by integrating RNA sequence, secondary structure, and 3-dimensional (3D) atomic arrangement information to identify nucleotide level binding residues, most of them neglect the molecular surface that directly contacts small-molecule compounds and are highly dependent on accurate three-dimensional structures. Here, we present MultiRSF, a multimodal deep learning framework that integrates molecular surface fingerprints, contextual sequence embeddings from a pretrained RNA language model and dot bracket secondary structure encoding to predict ligand binding sites on RNA. MultiRSF fuses these features through a hierarchical Transformer encoder and demonstrates robust predictive performance on both ligand-free RNA structures (apo RNA) and ligand-bound RNA structures (holo RNA). On independent benchmark sets TE18 and APO8, MultiRSF outperforms state-of-the-art methods, achieving precision of 0.765/0.458, recall of 0.698/0.286, and Matthews correlation coefficient (MCC) of 0.535/0.279. Case studies on a ligand-induced riboswitch conformational change and an NMR conformational ensemble further illustrate the model’s robustness to moderate RNA flexibility. In conclusion, MultiRSF provides an accurate and generalizable tool for nucleotide resolution binding sites prediction, with potential to accelerate early-stage RNA-targeted drug discovery.
Identifying enzymes capable of catalyzing specific chemical transformations across large sequence databases remains a major challenge in biocatalyst discovery. Conventional fingerprint-based methods capture global molecular structure but fail to represent bond-breaking and bond-forming events, limiting generalization to structurally novel reactions. We introduce a dual-track evaluation framework to distinguish true generalization from memorization, assessing retrieval on structurally isolated reactions (n = 50) within a 63,259-sequence enzyme pool. The strongest fingerprint baseline achieves R@10 = 0.020. To address this limitation, we develop GATv2-ECR, a heterogeneous dual-tower model integrating reaction-center graph encoding, a frozen ESM-2 sequence encoder, contrastive learning, and EC-aware soft reranking. GATv2-ECR achieves R@10 = 0.160 on isolated queries and R@10 = 0.308 on an out-of-distribution subset (n = 39), capturing mechanistically relevant features and supporting generalizable enzyme retrieval under open-world conditions.
Abstract ThermoLIB is a Python/Cython library designed to be used as a postprocessing tool for constructing free-energy surfaces, including adequate error estimation from the output of molecular simulations, transforming them between different collective variables (CVs), and extracting thermodynamic and kinetic information. ThermoLIB is available for download on GitHub and comes with extended documentation as well as many tutorials. The implementation is based on the theory of maximum-likelihood estimators for robust error estimation. The free-energy surfaces can be transformed or projected a posteriori to a larger or lower CV space by constructing conditional probabilities from the simulation results. Finally, useful CV-independent thermodynamic and kinetic properties, such as the rate constant, can be readily obtained, together with their uncertainty estimates. We briefly illustrate the capabilities of ThermoLIB by means of tutorials and case studies.
Abstract Bayesian optimization (BO) is widely used for reaction optimization, but under small-data conditions, regression of absolute objective values may not align with the practical goal of prioritizing promising experiments. Here, we compared ranking-based BO (RankBO) combined with Thompson sampling (Rank-TS) with regression-based BO using two reaction benchmarks. Rank-TS showed a clear advantage on Direct Pd-catalyzed arylation in optimization performance, global ranking quality, and recovery of high-yielding conditions. For Suzuki–Miyaura coupling, which comprised 12 substrate combinations, its improvement was small in the pooled analysis. However, in the equal-weight macro analysis across the 12 substrate-defined search spaces, Rank-TS showed higher ranking quality and faster high-yield-condition recovery. These results indicate that Rank-TS is particularly useful for prioritizing reaction conditions within a fixed-substrate combination under small-data conditions, and that its relative performance depends on how the chemical search space and ranking task are defined.
Abstract Accurate prediction of equilibrium partition coefficients of organic compounds in biological and environmental media is critical for evaluating their environmental fate and bioaccumulation potential, as well as for guiding drug design and toxicology studies. Linear solvation energy relationships using Abraham solute descriptors (ASD-LSERs) have achieved remarkable success in characterizing such complex biological and environmental partitioning systems. However, the limited availability of high-quality experimental descriptors, particularly for structurally complex compounds, hinders their practical application and further constrains the accuracy of descriptor predictive models. Here, we propose a linear solvation energy relationship (4SD-LSER) using descriptors derived from logarithmic n-hexadecane–air, n-octanol–air, and water–air partition coefficients, along with the topological McGowan molar volume. To evaluate its performance, 1,849 experimental partition coefficients for 779 neutral compounds across 12 biologically and environmentally relevant systems were compiled. The 4SD-LSER was calibrated and exhibited good descriptive power for these systems. Remarkably, when combined with appropriate fragment-based or machine learning-based descriptor prediction approaches, the 4SD-LSER achieved prediction errors largely within ±0.5 log units for structurally simple compounds and within ±1.0 log unit for more complex compounds (e.g., pesticides, pharmaceuticals, and flame retardants), exhibiting state-of-the-art accuracy, especially for complex compounds. This study demonstrates that models originally developed for well-characterized solvent systems to predict partition coefficients or solvation free energies can be readily extended to biological and environmental systems via LSER. Its performance is poised to improve further with advances in theoretical and machine-learning approaches.
Molecular representations that capture both chemical information and property-relevant differences are essential for reliable molecular property prediction in drug discovery and toxicity assessments. Fragment-level SMILES representation learning is effective at modeling substructure semantics, but structural reconstruction objectives alone do not explicitly encode the descriptor-derived physicochemical information needed for downstream prediction, especially in small data and scaffold-split settings. To address this limitation, we propose PKP-DiffMol, a physicochemical knowledge-prompt encoding and latent diffusion framework built upon the pretrained SMI-EDITOR encoder. PKP-DiffMol improves fragment-level molecular representations by using numerical and semantic physicochemical knowledge. The numerical physicochemical prior branch encodes RDKit descriptors as continuous quantitative priors, while the physicochemical knowledge prompt encoder transforms the same descriptors into semantic physicochemical representations by using a Transformer-based text encoder. The two complementary views are then integrated through hierarchical physicochemical knowledge fusion to construct a chemically enriched, fused latent space. In this space, a quality-controlled conditional latent diffusion module learns label-conditioned distributions, generates synthetic molecular latent representations, and retains reliable samples through a Mahalanobis distance quality gate. Experiments on seven MoleculeNet classification datasets under scaffold splitting show that PKP-DiffMol increases the mean ROC-AUC from 77.80% for the SMI-EDITOR backbone to 80.53% and obtains the best result on five of the seven datasets. The largest gains are observed on BACE (+6.91%), MUV (+4.29%), and SIDER (+3.92%), showing strong predictive performance across bioactivity prediction, virtual screening, and adverse drug reaction prediction tasks. Further analyses indicate that numerical and semantic physicochemical knowledge improves class separability and the correspondence between representation distances and RDKit descriptor distances, while Mahalanobis-filtered label-conditioned latent augmentation supports prediction-head refinement.
Abstract The fast multipole method (FMM) for computing Coulombic forces in molecular dynamics (MD) simulations offers benefits for a variety of systems, including nonperiodic configurations such as molecules in vacuum and large-scale systems where many other approaches encounter limitations. Because FMM parameters govern both simulation accuracy and computational performance, and because the optimal solution is highly system-dependent, we have developed FMM-AutoOpt for the automated benchmarking of FMM parameters with GROMACS. Relying solely on standard Python libraries and GROMACS FMM (Kohnke, B.; et al.J. Chem. Theory Comput.2020, 16, 6938–6949.), the tool is versatile and flexible, facilitating rapid parameter screening to ensure optimal performance and high simulation accuracy.
Abstract Recently, a new category of machine learning approaches for tabular data has emerged: tabular foundation models (TFM), based on in-context learning. A TFM is a neural network (usually a transformer) pretrained primarily on synthetic data. Its input is an entire data set: features and labels for training records, along with features for the target records. This study comprehensively benchmarked these methods for predicting the properties of small organic molecules. Several TFM were compared with multiple machine learning methods across 11 data sets (regression, random and structure-aware splits, up to 10,000 molecules each). The evaluation showed that TFM consistently outperform XGBoost, CatBoost, multilayer perceptrons, and other descriptor-based methods, even with careful hyperparameter selection for the baselines. TFM also demonstrate accuracy on par with or better than graph-based methods, including those pretrained on chemical data. Uni-Mol2, a pretrained deep neural network operating on 3D atomic coordinates, slightly outperforms TFM in some experiments. However, this comparison deliberately disfavors TFM, as they do not use any chemical pretraining and rely on a minimalistic set of 2D molecular descriptors without feature engineering. Some further improvement in results for relatively large data sets and random splits is achieved using retrieval: for each test molecule, the 500 closest neighbors (by Tanimoto similarity) are selected from the training set, and TFM inference is performed on this local subset. Overall, TFM (particularly the TabPFN family) are highly promising for predicting the properties of small molecules.
Abstract G-quadruplexes (G4) are noncanonical nucleic acid secondary structures formed by guanine-rich sequences that can form higher-order multimers under appropriate conditions and represent attractive targets for multivalent molecular recognition. Here, we propose a rational ligand design strategy based on multivalency, in which ligand molecules are covalently linked into polymer-like macromolecules. Using coarse-grained molecular dynamics, we show that these “polyligands” intercalate G4 multimers more efficiently than their monomeric counterparts, achieving strong stabilization through cooperative effects, even when individual ligands exhibit low intrinsic binding affinity. Polyligand intercalation induces structural changes in G4 multimers, modifying their size and stiffness. Using a quantitative polymer theory framework, we connect these structural changes with the thermal stabilization of G4 multimers. Polyligands represent a general design principle for G4-targeting molecular stabilizers, offering enhanced G4 stabilization via programmable linker length and backbone rigidity.
Abstract Predicting the impact of single-point mutations on protein thermodynamic stability is crucial for protein engineering of therapeutic and industrial applications. By effectively capturing the three-dimensional structural information of proteins and the spatial physical environment of each residue, the inverse folding models (IFMs) upon fine-tuning, such as ThermoMPNN, achieved state-of-the-art performance in predicting thermostability changes in proteins caused by mutations. However, IFMs are limited in their capacity to capture protein deep evolutionary information, whereas protein language models (pLMs) excel. Here, we present SA-MPNN, a lightweight, end-to-end hybrid framework that dynamically integrates the protein sequence representations from a protein language model (ESM2) into the ThermoMPNN architecture to improve protein stability prediction by combining evolutionary representations with geometric structural embeddings. By evaluating various feature fusion strategies, we selected a self-attention-based integration mechanism to effectively combine the two modalities. Trained on the large-scale Megascale data set, SA-MPNN achieved modest but consistent gains over ThermoMPNN on various benchmark data sets, with particularly noticeable improvements in several correlation analysis and screening-oriented evaluations. Finally, wet-lab validation was performed on the top-ranking variants of Acetivibrio thermocellusβ-glucosidase (AtBgl1A) as a case study. The experimental results demonstrated that multiple designed mutants exhibited improved thermostability, and the optimal variant, GC20, achieved a melting temperature (Tm) of 76.98 °C, representing a 5.97 °C increase over the wild-type, thereby supporting the practical applicability of SA-MPNN in protein engineering.
Abstract Ephrin receptors (Ephs) are receptor tyrosine kinases that regulate cellular growth, differentiation, and motility. EphA2, often overexpressed in cancer, is notable for its ligand-independent activation, which drives pro-oncogenic signaling distinct from the canonical, ligand-dependent pathway that restricts cell movement. While ligand binding induces extracellular clustering, kinase activation depends on dimerization within the transmembrane (TM) region. EphA1 and EphA2 differ substantially in their function, likely due to differences in their TM, but more so in their juxtamembrane (JM), and membrane-proximal fibronectin type III (FN1/FN2) domains. How these latter two regions modulate dimerization has not yet been investigated. To address this, we performed extensive coarse-grained (CG) simulations using Martini3 in an anionic POPC/PS/PIP2 model of the plasma membrane. Both receptors formed stable TM dimers, though EphA1 favored a symmetric AXXXGXXXG-centered interface, whereas EphA2 favored an alternate leucine zipper interface. All protein constructs sampled multiple configurations, reflecting substantial intrinsic variability. AlphaFold3 does not yield reliable predictions and cannot account for the effects of the lipid bilayer. In the CG simulations, basic residues in the JM region remained membrane-bound, and the EphA2 FN domain displayed sustained PIP2 interactions, consistent with previous observations. Notably, constructs with the FN2 domain alone (a fragment relevant to Alzheimer's Disease) restricted TM association in both receptors, whereas inclusion of the second FN1 domain restored dimerization but produced receptor-specific extracellular interfaces. These differences arise from distinct FN1-FN2 linker flexibilities, a different level of sequence conservation of EphA1 compared to several other Ephs and FN-domain membrane contacts, which together shape TM geometry and lipid engagement. Our results predict how TM and TM-proximal elements cooperatively tune dimerization in Eph receptors. This work offers a molecular rationale for the divergent activation behaviors of EphA1 and EphA2 and provides testable models relevant to cancer and neurodegeneration biology as well as Eph-driven signaling.
Abstract Protein–protein interactions (PPIs) regulate essential cellular processes and represent an important class of therapeutic targets; however, discovering effective modulators of PPIs remains a formidable challenge. Although deep learning approaches have been widely explored for PPI modulator discovery, many rely on simplified representations that obscure interchain boundary information and fine-grained PPI–modulator interaction (PPIMI) patterns, limiting their robustness under distribution shifts. To address this challenge, we introduce TvTPPIMI, a framework that leverages learnable boundary tokens to encode partner-aware boundary information and models PPIMI at atom–residue resolution. In a case study targeting the AURKA–TPX2 interaction, TvTPPIMI prioritized putative modulatory candidates from a small-molecule screening library. Structure-based docking, attention analysis, multireplica molecular dynamics simulations, MM/GBSA binding free-energy estimation, and noncovalent interaction analyses provided post hoc physical support for the stable AURKA binding of selected candidates, highlighting CE02–6266 as the most favorable compound among the tested hits. Together, these results suggest that TvTPPIMI provides a generalizable computational framework with coarse-grained, attention-based interpretive cues for PPIMI prediction and can be integrated with structure- and dynamics-based analyses to support PPI modulator discovery.
Abstract Specific recognition between T-cell receptors (TCRs) and peptide–major histocompatibility complexes (pMHCs) is central to adaptive immunity, yet accurate prediction of TCR–pMHC specificity remains challenging. Existing models mainly rely on sequence features or isolated molecular structures, limiting their ability to capture interface-level determinants within the ternary recognition complex. Here, we constructed the multimodal TCR–pMHC ternary complex (MM-TCR) data set, integrating paired TCR–pMHC sequences, V/J gene annotations, and modeled TCR–pMHC complex structures refined by short molecular dynamics-based relaxation. Based on MM-TCR, we developed TCRspec, an interpretable multimodal framework combining sequence embeddings, gene-usage features, and complex-level structural representations. Under a stringent CD-HIT TCR-cluster-disjoint split, TCRspec achieved an average AUROC of 0.896 and AUPRC of 0.882 across seven antigen-specific test data sets, outperforming representative baseline models. Cross-validation and ablation analyses confirmed the contribution of ternary complex structural information and MD-refined structures. In independent OOD peptide–TCR systems, TCRspec retained discriminative performance and identified model-inferred peptide positions associated with TCR recognition, providing a structure-informed framework for TCR specificity prediction.
Abstract This review addresses a critical need in the emerging intersection of digital chemistry and material properties by critically evaluating the current capabilities and future prospects of applying machine learning (ML) to ionic liquids (ILs) for thermal energy storage (TES) applications. We examine available data sets for key thermophysical properties relevant to TES, including thermal conductivity, heat capacity, density, viscosity, surface tension, melting point, enthalpy, CO2 solubility, and toxicity. Particular emphasis is placed on chemical space coverage, biases, and practical limitations. We compare different modeling approaches, spanning from conventional QSPR/QSAR and group-contribution methods to contemporary ML approaches, including kernel models, ensemble methods, graph neural networks (GNNs), and deep learning architectures. The impact of molecular representations and descriptor choices and validation strategies in determining model robustness and transferability is examined. To structure this assessment, we introduce a three-tier property-selection framework for IL-based TES and apply a six-criterion evaluation checklist consistently across all nine reviewed properties. The analysis is supported by literature-summary tables that consolidate data sets, chemical space coverage, modeling approaches, and validation practices across different property domains. Finally, we outline recommendations for data set standardization, open sharing of data and code, and the establishment of benchmark validation protocols to improve reproducibility and accelerate the discovery of ILs for TES applications.
Abstract Molecular property prediction is a key task in AI-driven drug discovery, yet the prevalence and impact of label imbalances in molecular property regression remain poorly understood. Through a systematic benchmark of widely used molecular property data sets, we show that target values are often highly imbalanced and that prediction errors are consistently concentrated in sparsely represented regions of the label space. We further evaluate representative imbalance-learning approaches developed for general regression tasks and find that their effectiveness on molecular data sets is limited. To address this challenge, we propose interval-aware mixture-of-experts (IA-MoE), a plug-in framework that partitions the continuous target space into intervals and promotes expert specialization across different regions of the label distribution. IA-MoE can be seamlessly integrated with diverse molecular encoders without modifying their underlying architectures. Experiments on five molecular property benchmarks and four graph neural network backbones demonstrate that IA-MoE consistently improves the predictive performance in underrepresented regions while maintaining or improving overall accuracy. Our findings establish label imbalance as an important challenge in molecular property regression and highlight interval-aware expert learning as an effective strategy for addressing this challenge.
Abstract Accurate prediction of protein–ligand binding poses is essential for structure-based drug discovery. In recent years, computational approaches, particularly molecular docking, have generated large numbers of candidate ligand poses. However, existing scoring functions remain limited in their ability to distinguish near-native poses from incorrect docking decoys. In this work, we propose UniPoseScore, a Graphormer3D-based framework for ligand pose scoring and refinement. The model simultaneously predicts pose RMSD and atomic displacement vectors to guide ligand pose refinement. Across four benchmark data sets, UniPoseScore showed consistently strong performance and better generalization, with a particularly clear advantage in RMSD correlation. Visualization of the learned embeddings further suggests that the model captures structural features correlated with pose accuracy. On the CASF-2016 benchmark, UniPoseScore achieves a refinement success rate of 74.3%, highlighting the effectiveness of UniPoseScore in improving ligand binding poses and its potential applications in drug discovery. The UniPoseScore is available at https://github.com/April-Zhangjq/UniPoseScore.
Abstract Fub7 is a rare pyridoxal 5′-phosphate (PLP)-dependent enzyme that can subsequently catalyze the γ-elimination of O-acetyl-l-homoserine (OAH) and Michael addition with n-valeraldehyde (NVA) to synthesize 5-alkyl-pipecolic acid, which shows significant potential in synthetic chemistry. In this study, we constructed the computational models and performed MD and QM/MM calculations to illuminate the complete catalytic cycle of Fub7, which contains five reaction stages, including external aldimine formation, γ-elimination, Michael addition, reformation of the internal aldimine, and final cyclization and dehydration, of which γ-elimination and Michael addition correspond to relatively high energy barriers. In addition, according to our test calculations, Fub7 can also catalyze the β-elimination. During the catalysis, the terminal amino of Lys211 exhibits significant conformational changes and was found to play multiple roles. On one hand, Lys211 forms internal aldimine with the cofactor PLP. On the other hand, Lys211 functions as a catalytic base/acid or as a bridge to mediate a series of proton transfer in the γ-elimination and Michael addition. Tyr109 only participates in the Michael addition. The deprotonated Tyr109 acts as a base to abstract the α-proton of NVA to generate the carbanion nucleophile, which is the key for Michael addition. These findings may provide useful information for understanding the catalysis of PLP-dependent enzymes and designing biocatalysts for γ-substitution.
Abstract Mimotopes selected by phage display have been used for antibody epitope mapping for nearly four decades, but residue-level recovery of functional epitopes has remained unreliable. This is not a problem of algorithms but of assumptions: mimotopes recapitulate native epitope recognition through compensatory interaction networks that preserve binding energetics, without conserving sequence or surface geometry. Predictors trained on sequence similarity or static interface geometry therefore miss the residues that drive recognition. We present a physics-based workflow that refines antibody-bound mimotope conformations by microsecond-scale molecular dynamics and maps the resulting structures onto the cognate antigen using Folddisco. Across four antibody–antigen systems with crystallographically defined epitopes, the workflow reaches 60 to 100% precision in epitope localization and recovers 100% of mutagenesis- or structurally validated functional hotspots, compared with at most 25% for three widely used benchmarks (EpiSearch, ClusPro, SEPPA-mAb). For the clinically relevant ChiLob 7/4 system, it further identifies a minimal tetrapeptide (PWVP) sufficient for antibody recognition, validated by Western blot and immunoprecipitation. By grounding epitope prediction in experimentally selected binding information and energy-resolved conformational refinement, the workflow connects phage display to mechanistic structural insight.
Abstract The detailed mechanism of halohydrin dehalogenase (HHDH)-catalyzed ring-expansion reaction between spiro-epoxides and the nucleophile OCN– to form spiro-oxazolidinones was investigated using molecular docking, molecular dynamics (MD) simulations, and quantum mechanics/molecular mechanics (QM/MM) calculations. Molecular docking and molecular mechanics/Poisson–Boltzmann surface area (MM-PBSA) results revealed that the R-configurational substrate exhibited significantly superior binding affinity compared to the S-configurational one, validating the stable binding mode within the enzyme′s active pocket. The fundamental pathway initiates with the nucleophilic attack of OCN– accompanied by proton transfer, which is followed by ring closure coupled with a proton transfer to form a five-membered ring. A total of four possible selective nucleophilic attack pathways were evaluated, and the most energetically favorable pathway associated with the R-configurational product has an energy barrier of 17.1 kcal/mol, which is close to the experimental value of 18.1 kcal/mol derived from kinetic parameters. This study provides valuable theoretical insights for the future engineering of HHDHs for chiral spiro-oxazolidinone synthesis.