
We describe a fully optimized non-equilibrium (NEQ) alchemical workflow for predicting relative binding free energies using the Q molecular dynamics engine. The method is implemented as a function within QligFEP...
Materials discovery has been a recurring testbed for advances in machine learning: graph neural networks for property prediction, diffusion models for crystal generation, and universal interatomic potentials trained on millions...
DaHAC combines dual large language models with human expert adjudication to extract high-fidelity CRISPR biosensing knowledge from literature, building a validated dataset and benchmark for evaluating LLM reasoning in biosensing.
Collective variables (CVs) are low-dimensional projections of molecular configuration space that serve a dual purpose: they provide mechanistic interpretability by distilling complex transformations into comprehensible reaction coordinates, and they underpin the majority of enhanced sampling methods by defining the directions along which free-energy barriers are overcome. The discovery of suitable CVs has evolved from reliance on chemical intuition, for instance by selecting distances, angles, or dihedrals by hand, the systematic linear approaches, such as principal component analysis and time-lagged independent component analysis, to nonlinear manifold-learning techniques including diffusion maps and variational autoencoders. Deep-learning methods have further expanded this landscape: committor-based neural networks approximate the optimal reaction coordinate from trajectory data, discriminant and variational models learn CVs tied to metastable-state separation or slow kinetics, and equivariant graph neural networks construct symmetry-preserving representations directly from atomic coordinates. In parallel, generative models-normalizing flows, diffusion-based samplers, and learned transfer operators-have begun to bypass explicit dimensionality reduction altogether, learning equilibrium distributions or dynamical propagators in the full configurational space. Yet, these models encode latent structure from which CVs can be extracted a posteriori for mechanistic interpretation. In this Perspective, we trace the arc of CV discovery from intuition-driven heuristics to modern data-driven and generative frameworks, critically assess the strengths and limitations of each class of methods, and outline how the convergence of machine-learned potentials, automated CV learning, generative sampling, and causal interpretability is giving rise to integrated workflows that will reshape predictive molecular simulation across biomolecular, catalytic, and materials systems.
We review molecular and polymer representations for machine learning, from descriptors and fingerprints to learned embeddings, highlighting how hierarchy, stochasticity, and data limitations shape polymer informatics.
Data-driven materials discovery requires large-scale experimental datasets, yet most of the information remains trapped in unstructured literature. Existing extraction efforts often focus on a limited set of features and have...
Property-conditioned molecular generation enables precise targeting of enthalpy of formation, producing CHON energetic molecules with properties close to specified EoF targets.
Chemical imputation models provide a flexible strategy for estimating missing values in sparse chemical databases.
Homolytic bond dissociation energy (BDE) governs radical formation and bond-breaking thermodynamics, yet in molecular design it is typically evaluated only after candidate structures are proposed. Here we treat BDE as...
This review provides an overview of DEL screening challenges, computational strategies for de-noising enrichment counts, DEL-related ML challenges, and hit prioritization and identification, with a particular focus on machine learning approaches.
An LLM-assisted, deterministically post-processed pipeline translates literature-derived χDL procedures into Chemspeed execution programs with row-level provenance, validated by robotic synthesis (84.4% yield) and input-robustness benchmarking.
A physics-guided machine-learning pipeline converts a measured cell deformation index into fabrication-ready DLD pillar geometries in under a minute, replacing weeks of microfluidic trial-and-error with an open web tool.
Influenza A virus (IAV) remains a persistent global health threat. Rapid genetic evolution of IAV strains resistant to known drugs underscores the urgent need to identify novel antiviral scaffolds. In this context, natural products provide chemically diverse scaffolds. However, for many reported anti-IAV phytochemicals, the precise viral targets and binding sites remain unclear. Here, we applied an integrated AI-assisted computational workflow to identify druggable IAV targets for six phytochemicals, including diphylin (DP), justicidins A–C (JA, JB, and JC), and tabamide derivatives (TA and TA25) across clinically dominant H1N1 and H3N2 subtypes. DiffDock was used for protein-wide binding-site identification, followed by focused docking with five independent AutoDock Vina replicates per ligand–target pair and GNINA CNN rescoring to jointly assess affinity and pose reliability. Among the twelve viral proteins examined (HA, NA, RdRp, M2, NP, and NS1 for each subtype), RdRp and NP showed the most consistent and favorable interaction profiles, while HA and NS1 were comparatively weak targets. Because an experimental H1N1 RdRp structure lacked a large PB2 region, an AlphaFold3-predicted model (AF3P) was incorporated to enable complete PB2-domain assessment. Lead complexes were identified and processed for 100 ns molecular dynamics simulations, followed by binding free energy calculations using MM-GBSA and MM-PBSA methods. The combined results of MD simulations, interaction patterns within the PB2 cap-binding pocket, and free-energy ranking identified DP and JB as the most promising broad-spectrum candidates. These computational findings are encouraging for further optimization and experimental validation of lignan scaffolds targeting the conserved PB2 cap-binding domain of IAV RdRp.
The high demand for developing next generation polymers has led to the use of data-driven approaches, such as inverse design, to exploit the chemical space for polymer design. Existing optimization-based approaches for the inverse design of polymers typically either utilize a single objective function or a Pareto front-based multi-objective optimization strategy. These methods tend to trade off properties of interest, achieving neither with precision, and are computationally intensive. In this paper, we argue that target properties can often be partitioned into primary properties, which must be achieved with high accuracy, and secondary properties that are of interest to achieve, but not at the expense of achieving worse outcomes for the primary properties. We introduce an algorithm for experimental design of polymers that adheres to this principle when identifying solutions in regions of the chemical design space. The algorithm performs well, accurately identifying desired target values even in high dimensions, with solutions that are in line with established domain knowledge.
Quantum machine learning (QML) holds significant promise for chemistry applications, yet practical implementation faces fundamental challenges, particularly the barren plateau phenomenon that prevents effective training of deep quantum circuits. Here, we present a novel quantum neural network architecture that successfully overcomes this limitation through sign-alternated angle encoding (SAE) combined with identity-block (IB) initialization and data re-uploading (DRU). While quantum data re-uploading has been shown to enhance expressibility in single-qubit systems, we demonstrate that it exhibits limitations in multi-qubit circuits, causing performance degradation and increased sensitivity to initialization. Our proposed DRU + IB + SAE architecture addresses these issues by applying sign alternation to even-numbered encoding layers, enabling successful training of deep 12-qubit quantum circuits while retaining strong gradient signals throughout optimization. We validate our approach on a real-world chemical problem: binary classification of molecular toxicity using 10 448 molecules represented in SMILES format. The proposed architecture achieves the highest performance in accuracy (0.812) and F1 score (0.767) compared to baseline quantum neural networks and comparable QNN-less models, while also demonstrating improved stability across trials with reduced variance across different initializations. Through comprehensive gradient analysis measuring root mean variance (RMV), we show that our method effectively mitigates exponential gradient decay even at circuit depths where conventional approaches fail. This work bridges the gap between theoretical quantum computing advances and practical chemical applications, demonstrating that carefully designed quantum architectures can provide advantages for real-world chemical problems even with classical data in the current noisy intermediate-scale quantum (NISQ) era.
Artificial intelligence (AI) is rapidly transforming the field of biosensing, enabling unprecedented advances in precision oncology and multi-omics integration. Biosensors have long provided sensitive platforms for detecting circulating tumor cells (CTCs), genomic mutations, proteomic signatures, and metabolic alterations, but their translational impact has been limited by challenges in data complexity and clinical interpretation. Recent breakthroughs in machine and deep learning now allow biosensors to move beyond signal acquisition toward intelligent data processing, predictive modeling, and real-time clinical decision support. Computational chemistry and AI-driven algorithms are increasingly applied to optimize sensor–analyte interactions, enhance detection specificity, and integrate heterogeneous omics datasets into unified diagnostic frameworks. This convergence of biosensing technologies with computational modeling and precision medicine strategies is reshaping cancer diagnostics, enabling early detection, patient stratification, and therapy monitoring with unprecedented accuracy. By synthesizing recent landmark studies (2023–2025) alongside foundational work in biomarker profiling and CTC detection, this review highlights the transformative potential of AI-enabled biosensors as smart platforms for oncology and multi-omics. The integration of biosensing with AI not only addresses current limitations in reproducibility and scalability but also paves the way for clinically validated personalized diagnostic systems that can accelerate the adoption of precision oncology worldwide.
AI-assisted prediction of chemical reactions has garnered significant attention and research over the past several years, particularly in the fields of synthetic and medicinal chemistry. To improve the prediction accuracy, many studies have incorporated data augmentation techniques for enhancing the generalizability and flexibility of models. However, artificial reaction generations have been largely limited to SMILES enumeration and minor modifications around the reaction center, resulting in conservative transformations. To generate more diverse artificial reactions, we developed several augmentation modules using a cluster-based strategy and a reaction site-centered buffer zone technique. Notably, our augmentation modules provided reliable data and thus generally led to improved performance in retrosynthesis prediction. In addition, the combined use of the modules provided superior results, which suggests the potential for broader application of our approach in reaction prediction tasks and beyond. Furthermore, we have confidence that this work will open up opportunities to extend the usability of such methods toward valuable generative models for chemical reactions and other downstream tasks in computational chemistry and drug discovery.
Effective drug storage and delivery systems remain a major challenge for complex diseases such as cancer, where maximizing therapeutic efficacy while minimizing side effects is critical. Metal–organic frameworks (MOFs) offer a highly tunable platform for drug encapsulation, yet systematic exploration across the vast MOF design space is still limited. In this work, we presented an integrated, data-driven computational framework that combines density functional theory (DFT) calculations, molecular simulations, and machine learning (ML) to screen 90 653 synthesized and hypothetical MOFs, the largest dataset studied to date for drug storage applications. Focusing on 5-fluorouracil (5-FU) and methotrexate (MTX), we showed that MOFs can achieve exceptionally high drug adsorption capacities (up to 18.6 and 17.1 g g−1, respectively), significantly exceeding experimentally reported values for a limited number of MOFs. ML models trained on simulation data demonstrated strong agreement with available experimental data, enabling rapid and reliable prediction across diverse chemistries. Structure–property relationships were elucidated through detailed analysis of MOF–drug interactions, including multi-drug adsorption behavior, while molecular dynamics (MD) simulations were used to provide molecular-level insight into drug mobility within MOF pores. Our work establishes a scalable and transferable digital discovery workflow for accelerating the design of MOFs for drug storage and delivery.
Herein, we present a theoretical in silico approach for evaluating the antioxidant properties of known and potential compounds using conventional DFT calculations. The method is based on the total potential...
Steric descriptors are central to understanding and comparing chiral phosphoric acid (CPA) catalysts, yet practical use of these parameters often remains limited by manual structure inspection and the absence of...