
ABSTRACT Equivariant graph neural networks (EGNNs) are becoming the geometric infrastructure of 3D molecular generation for AI‐aided drug discovery. By enforcing E(n), E(3), or SE(3) equivariance, they separate physical molecular structure from arbitrary coordinate‐frame conventions and ensure that predicted coordinates, denoising directions, and velocity fields transform consistently with molecular geometry. This symmetry‐aware design supports direct modeling of conformations and protein–ligand spatial relationships while remaining compatible with diffusion, flow‐based, autoregressive, and hybrid generative frameworks. Consequently, EGNNs enable more geometrically consistent generation and facilitate controllable design conditioned on binding pockets, pharmacophores, fragments, scaffolds, reference ligands, or molecular properties. Their benefits, however, are bounded by what equivariance encodes. Coordinate‐frame consistency does not itself enforce valid bonds, valence, stereochemistry, topology–geometry agreement, synthesizability, or biological activity. Moreover, practical models must reconcile discrete chemical variables with continuous geometric dynamics, represent long‐range interactions without prohibitive computational cost, and account for protein flexibility, solvent, metal ions, and induced fit. Performance is further constrained by biased docking or property predictors, heterogeneous benchmarks, and limited prospective validation. This review synthesizes how EGNNs function across major generative paradigms and argues that future drug‐oriented systems should combine joint 2D–3D graph generation, physically credible interaction modeling, synthesis and ADMET constraints, multi‐objective optimization, and closed‐loop experimental feedback. EGNNs should therefore be viewed not as a complete chemical solution, but as the symmetry‐aware foundation upon which more testable and biologically relevant molecular design systems can be built.
ABSTRACT Spatial confinement radically changes the collective behavior of amphiphilic molecules, generating self‐assembled entities that significantly exceed bulk systems in morphological and topological complexity. Under geometric persuasion (and a modest nudge from entropy), these molecules orchestrate unexpected architectures: vesicle‐in‐vesicle nesting, fenestrated structures, and convoluted hydrophilic/hydrophobic chambers with curious properties for intracellular transport. Remarkably, confined systems echo the cell's spatially restricted nature, depicting fusion, fission, and engulfment events, suggesting that confinement does not merely impose boundary constraints but natural design principles: tune the cage, and the molecules will improvise a new choreography. Within nanosciences, recognizing confinement as a controllable parameter extends the traditional concept of self‐assembly from a passive chemical consequence into an engineerable route toward hierarchical 3D structures with programmable form and function, each one a little kingdom with its own rules for curvature and asymmetry. This article is categorized under: Structure and Mechanism > Computational Biochemistry and Biophysics Structure and Mechanism > Computational Materials Science Molecular and Statistical Mechanics > Molecular Dynamics and Monte‐Carlo Methods
ABSTRACT Artificial intelligence (AI) is reshaping drug discovery by accelerating timelines and reducing costs, yet its impact remains constrained by a persistent gap between computational promise and translational delivery. This gap stems from upstream preclinical failures, including weak target validation, biologically irrelevant models, and insufficient accountability for overstated methodological claims that contribute to late‐stage attrition. The Implementation, Methodology, Productivity, Assessment, Collaboration, Translation (IMPACT) framework addresses these root causes by establishing global standards that reinforce biological grounding, methodological credibility, and equitable collaboration. Implementation emphasizes Findable, Accessible, Interoperable, and Reusable (FAIR)‐compliant datasets, standardized vocabularies, and clear gradients of AI involvement from assisted to fully AI‐driven workflows. Methodology prioritizes reproducibility through model cards, containerized environments, and transparent reporting to support robust models. Productivity aligns AI efforts with urgent therapeutic priorities, including rare diseases, antimicrobial resistance, drug repurposing, and natural‐product discovery. Assessment promotes rigorous benchmarking, blind validation, and uncertainty quantification, drawing on the long‐established CASP model as a historical gold standard while critically examining emerging initiatives such as CACHE and Polaris Hub, which remain comparatively recent and evolving. Collaboration leverages federated learning, pre‐competitive consortia, and interdisciplinary teams integrating AI specialists with domain experts. Translation ensures outputs are explainable, clinically relevant, ethically aligned, and regulatory‐ready, consistent with emerging frameworks such as the FDA Draft Guidance on AI in Drug Development and the EU AI Act. By integrating technical standards with operational governance mechanisms, IMPACT provides a structured pathway toward transparent and translationally reliable AI‐driven drug discovery. This article is categorized under: Data Science > Artificial Intelligence/Machine Learning Software > Molecular Modeling Data Science > Chemoinformatics
ABSTRACT Simulating photoinduced charge transfer (CT) in the condensed phase is essential for understanding solar energy conversion. Traditional Marcus theory is limited by its assumption of a thermally equilibrated initial state, which is often invalid for photoinduced processes, where vertical excitation creates a nonequilibrium nuclear state. The subsequent structural relaxation requires a time‐dependent rate coefficient. This review focuses on Instantaneous Marcus Theory (IMT), an approach recently developed to capture these nonequilibrium effects. Derived as the classical limit of the nonequilibrium Fermi's golden rule (NE‐FGR), IMT provides a practical, Marcus‐like expression for the time‐dependent rate based on the dynamical average and variance of the donor‐acceptor energy gap. While the direct evaluation of IMT requires computationally expensive nonequilibrium molecular dynamics, the nonlinear‐response (NLR) formulation reformulates the theory in terms of efficient equilibrium molecular dynamics simulations. This framework has been extended to multistate systems, allowing the simulation of complex reaction networks through a set of coupled Pauli's master equations. We highlight the application of these methods to the carotenoid‐porphyrin‐fullerene molecular triad, a prototypical organic photovoltaic system, dissolved in organic solvent. For this system, IMT correctly predicts a transient enhancement of the CT rate by over an order of magnitude, a nonequilibrium effect missed by Marcus theory. The population dynamics from multistate IMT are in excellent agreement with results from all‐atom nonadiabatic semiclassical mapping dynamics and quantum NE‐FGR calculations. This work establishes the multistate NLR‐IMT method as a reliable and cost‐effective tool for simulating photoinduced CT dynamics in realistic condensed‐phase systems. This article is categorized under: Theoretical and Physical Chemistry > Reaction Dynamics and Kinetics Structure and Mechanism > Reaction Mechanisms and Catalysis Software > Simulation Methods
ABSTRACT The rapid evolution of machine learning (ML) has advanced materials discovery, providing tools to explore, predict, and design materials with tailored properties. Here we present an overview of emerging ML tools for data‐driven materials innovation, including data curation, feature engineering, model development, interpretability, and inverse design. We highlight high‐throughput material databases in providing large‐scale, DFT‐computed datasets, and discuss the importance of descriptor libraries that encode compositional and structural information into machine‐readable inputs for model development. Advances in ML architectures, ranging from classical algorithms to graph neural networks, are discussed for their ability to capture complex structure–property relationships. Particular emphasis is given to inverse design frameworks using generative models and optimization strategies to enable property‐targeted materials generation. We further explore interpretability and uncertainty quantification techniques that are important for bridging ML predictions with experimental validation. Automation platforms are described as tools for closed‐loop, high‐throughput discovery pipelines. We outline grand challenges, including data sparsity, model generalizability, and experimental integration. Finally, we summarize future directions that include foundation models pre‐trained on broad, multimodal materials data; self‐supervised learning strategies to reduce dependence on labeled datasets; ML workflows that embed thermodynamic and symmetry constraints to enhance interpretability; and fully autonomous laboratories that couple ML guidance with robotic synthesis and real‐time feedback. This article is categorized under: Structure and Mechanism > Computational Materials Science Data Science > Artificial Intelligence/Machine Learning
ABSTRACT Functional theories reformulate the many‐electron problem by expressing electronic properties as functionals of reduced quantities, providing efficient alternatives to wave function‐based correlation methods. Kohn‐Sham density functional theory (KS‐DFT) and reduced density matrix functional theory (RDMFT) exemplify this philosophy but remain limited by their single‐determinant nature and numerical complexity, respectively. This review presents hierarchically correlated orbital functional theory (HCOFT), a unified framework developed to overcome these limitations. By extending orbitals into tunable hypercomplex spaces and deriving hierarchically correlated orbitals (HCOs) with fractional occupations through Clifford algebra, HCOFT establishes the corresponding variational foundation and a continuous dimensional hierarchy that spans KS‐DFT, RDMFT, and the intermediate 1‐HCOFT—a third formal functional theory featuring paired HCOs that naturally capture strong correlation while maintaining computational stability. Further advances, including the explicit‐by‐implicit scheme for stable occupation optimization, the coupled optimization strategy for accelerated convergence through simultaneous orbital and occupation updates, and the development of short‐range screened, occupation‐dependent orbital functionals for balanced treatment of dynamical and strong correlation, further strengthen the practical applicability of HCOFT. By integrating mathematical rigor, algorithmic efficiency, and a flexible platform for functional construction, HCOFT provides a systematically improvable foundation for electronic‐structure modeling and offers a promising pathway toward a versatile and unifying paradigm for accurate first‐principles calculations. This article is categorized under: Electronic Structure Theory > Ab Initio Electronic Structure Methods Electronic Structure Theory > Density Functional Theory
The capability of anticipating and mitigating drug toxicity represents one of the most persistent challenges in drug development. Despite rigorous preclinical evaluation, nearly one third of drug candidates fail during the clinical phases due to safety issues, in particular hepatotoxicity and cardiotoxicity. Routine in vitro and in vivo toxicology tests, while essential to define safety margins, are frequently associated with high false‐positive rates and poor translation to human outcomes. While investigative toxicology has improved mechanistic understanding of off‐target interactions and toxicokinetics, the gap between preclinical findings and clinical safety still challenges efficient drug development. To overcome these drawbacks, the field is developing and applying more integrative strategies that combine computational modeling, machine learning, multi‐omics technologies, and advanced in vitro systems. These approaches propose predictive pipelines likely able to identify chemical toxicology issues during early design stages and to characterize mechanisms of adverse events. In this review we provide a critical overview of emerging tools and strategies for drug toxicity prediction, evaluating their current impact, limitations, and translational potential to reduce safety‐related attrition and support the development of safer therapeutics. This article is categorized under: Data Science > Artificial Intelligence/Machine Learning Molecular and Statistical Mechanics > Molecular Mechanics
Molecular Dynamics (MD) has established itself as a pivotal computational tool across various scientific domains, including chemistry, biology, and materials science. Despite its widespread utility, MD faces inherent challenges, such as accuracy limitations, computational speed, and sampling efficiency. In recent years, machine learning, particularly deep learning, has seen significant advancements and is increasingly being integrated into MD processes. This review explores how deep learning can mitigate the issues associated with MD by addressing them from multiple angles. However, deep learning techniques introduce their own set of hurdles, including the need for extensive data, issues of interpretability, high computational costs, and concerns regarding transferability. Here, we discuss recent progress in the field of deep learning to overcome these obstacles. Ultimately, our goal is to demonstrate that, by leveraging the advancements made in both the MD and the machine learning community, deep learning has the potential to significantly enhance the capabilities of MD, paving the way to new scientific discovery. This article is categorized under:
eMap is a web based application for predicting electron/hole transfer pathways in proteins based on their crystal structures. The predictions are based on the Pathways model, where each hop between electron transfer active (ETA) moiety is described through a tunneling inspired penalty function. This is eloquently rendered in the framework of graph theory, where each ETA moietie is represented by a node, and the edge lengths are related to the penalty function. eMap 2.0 takes this one step further by constructing graphs for multiple proteins, and then finds common subgraphs using frequent subgraph mining (FSM), specifically gSpan. Lastly, eMap 2.0 utilizes sequence and structural similarity measures to analyze the frequent subgraph mining results. Here, we show how this powerful method has been successfully utilized to rapidly provide insight regarding conserved pathways within protein families, to identify structures with mutations within protein families, and to differentiate between active and inactive structures.
The cover image is basThe cover image is based on the article Dissipative Particle Dynamics Modeling in Polymer Science and Engineering by Sousa Javannikkhah et al., https://doi.org/10.1002/wcms.70018 image
Heterogeneous catalysis has a wide range of applications in chemical manufacturing and sustainable technologies. It uses solid catalysis to enable efficient chemical transformations. Traditional research on active sites and reaction mechanisms relies heavily on experiments and computational methods, such as density functional theory calculations. However, the volume of scientific literature and data is growing fast. This rapid growth has made it increasingly difficult to capture, process, and act on emerging insights systematically. Recently, large language models (LLMs) have emerged as powerful tools to support various stages in catalysis research. Their ability to understand and generate natural language helps them extract useful information from vast amounts of text, assist in catalyst design, aid in planning experiments, and clarify complex descriptors. In this advanced review, we first analyze recent progress in applying LLMs to heterogeneous catalysis, focusing on four key areas: literature mining and knowledge extraction, catalyst design and screening, experiment automation and workflow optimization, and the interpretation of high‐dimensional descriptors. We then highlight the challenges in this field despite these advances, most notably the need for domain‐specific fine‐tuning and the improvement of molecular representation. We conclude by discussing future opportunities for integrating LLMs with complementary machine learning approaches and expert‐in‐the‐loop systems, toward accelerating the rational discovery of next‐generation catalysts. This article is categorized under: Data Science > Artificial Intelligence/Machine Learning Data Science > Chemoinformatics Theoretical and Physical Chemistry > Reaction Dynamics and Kinetics
The ability to computationally predict changes in protein thermostability upon mutation is crucial for advancing protein design and engineering, with applications ranging from therapeutics to biocatalysis. This review provides a comprehensive overview of the significant challenges and diverse computational strategies for predicting protein stability and understanding epistatic interactions across protein variants. A primary obstacle to this goal is the scarcity of high‐quality, large‐scale thermodynamic datasets, which are often biased toward single‐point, destabilizing mutations and lack standardized experimental metrics. This limitation directly impacts the performance and generalizability of data‐driven methods, from early machine learning approaches to modern deep learning architectures such as ThermoMPNN and protein language models. Physics‐based approaches, such as those employing Rosetta and FoldX energy functions, offer valuable insights but are often limited by their reliance on static structures and oversimplified representations of the unfolded state. While molecular dynamics simulations can capture the critical role of protein flexibility and dynamics in thermostabilization, their computational cost restricts their application in high‐throughput screening. Accurately predicting the effects of multiple mutations is further complicated by epistasis, where nonadditive interactions can significantly alter stability and function. Overcoming these hurdles requires a synergistic approach, integrating AI‐driven predictions with physics‐based simulations and accurate conformational sampling methods. Promising future directions include the development of more comprehensive and unbiased datasets, and improved modeling of epistasis and the (un)folded states and their ensembles. Such advancements are essential for enhancing the reliability of thermostability predictions and navigating the complex stability–activity trade‐offs inherent in protein optimization and design. This article is categorized under: Structure and Mechanism > Computational Biochemistry and Biophysics Data Science > Artificial Intelligence/Machine Learning Molecular and Statistical Mechanics > Molecular Dynamics and Monte‐Carlo Methods
Explainable artificial intelligence (XAI) is increasingly essential in drug discovery, where interpretability and trust must accompany predictive accuracy. As deep learning models, particularly, deep neural networks (DNNs) and graph neural networks (GNNs), enhance molecular property prediction, de novo design, and toxicity estimation, transparent, mechanistically meaningful insights become critical. This article classifies major XAI strategies in computational molecular science, including gradient‐based attribution, perturbation analysis, surrogate modeling, counterfactual reasoning, and self‐explaining architectures. Molecular representations, such as fingerprints, SMILES, molecular graphs, and latent embeddings, are evaluated for their impact on explanation fidelity. An evaluation framework is outlined using metrics like fidelity, stability, completeness, sparsity, and usability, with emphasis on integration into drug discovery workflows. The discussion also highlights emerging directions, including neuro‐symbolic systems and physics‐informed networks that embed mechanistic constraints into statistical models. By aligning algorithmic transparency with pharmacological reasoning, XAI not only demystifies black‐box models but also supports scientific insight, regulatory compliance, and ethical AI deployment in pharmaceutical research. This article is categorized under: Data Science > Artificial Intelligence/Machine Learning Data Science > Chemoinformatics Structure and Mechanism > Computational Biochemistry and Biophysics
By bridging molecular‐level insights with macroscopic performance metrics, computational strategies are poised to transform how we design next‐generation 3D‐printable materials with enhanced precision, functionality, and sustainability. We present a critical overview examining the role of computational methods in advancing the design and application of 3D‐printable polymers. We cover key considerations—including solvation behavior, viscosity, gel point, mechanical properties, and polymer structure—as well as the design of new polymer functionalities. We highlight how a spectrum of physics‐based methods, ranging from quantum chemical to coarse‐grained simulations, can be leveraged to interrogate relevant polymer properties at multiple scales. In particular, we illustrate the growing impact of machine learning in accelerating polymer discovery and optimization. Such methods, whether applied independently or integrated into multi‐scale modeling frameworks, offer powerful tools for pre‐screening and selecting optimal formulations tailored to diverse 3D printing technologies and applications. Although challenges remain to integrate different approaches into workable prediction pipelines, the rate of advance and improvements in methods, data interoperability, and data quality, offer great promise of a ‘concept to print’ pipeline in the future. This article is categorized under: Structure and Mechanism > Computational Materials Science Data Science > Artificial Intelligence/Machine Learning Structure and Mechanism > Molecular Structures
Aptamers—short single‐stranded DNA or RNA—are the latest biomolecules to fall within reach of powerful structure‐prediction pipelines that blend bioinformatics, computational chemistry, and artificial intelligence. These tools now enable high‐throughput exploration of aptamer conformational landscapes, a prerequisite for rational design and optimization of their exceptional target affinity and specificity. Next‐generation sequencing has democratized library analysis, allowing any laboratory to handle millions of variants. Hybrid workflows currently offer the most reliable secondary and tertiary structure models, and explicit treatment of conformational flexibility is proving indispensable for mapping binding‐competent states. Yet every predictive tier—from classic free‐energy minimization to deep learning—still underrepresents chemically modified nucleotides, the very substitutions that grant therapeutic aptamers nuclease resistance and pharmacokinetic longevity. Capturing the structural and dynamical consequences of these modifications remains the key unsolved problem. Progress, therefore, hinges on two fronts: richer parameterization and training data that encompass modified bases, and tighter coupling of in silico screens with biophysical and structural validation. Bridging these gaps will convert the current wave of computational advances into clinically relevant aptamer‐based drugs ready to be delivered to the patients. This article is categorized under: Structure and Mechanism > Molecular Structures Data Science > Computer Algorithms and Programming Data Science > Artificial Intelligence/Machine Learning
Integrating materials representations into feature engineering by rational design plays a critical role in determining the capability and accuracy of material property prediction via machine learning (ML). There still exists a lack of comprehensive classification and multi-dimensional evaluation for many existing feature models that could guide model selection in applications and further development. This review systematically classifies feature construction methods for crystalline structures, emphasizing the coupling between chemical and structural information. We systematically discuss the geometric configurations, chemical attributes, and their intricate coupling mechanisms that can be leveraged for feature engineering. Furthermore, a comprehensive comparison is performed across multiple aspects including graph network representation, structural information embedding, chemistry-structure information coupling, local versus global characteristics, long-range versus short-range description, algorithm compatibility with kernel function method or deep neural network, data size requirements, computational complexity, and interpretability mechanisms, thereby highlighting key variations in existing feature models and improving the physical interpretability of predictive models. To illustrate the integration of multi-dimensional characteristics, the center-environment (CE) feature model is introduced based on the coupling between local chemical and structural information of physical core-shell structures. Within the CE model, the pre-attention mechanism reorients focus from intricate details within complex ML algorithms to explicit feature models that depict physical core-shell configurations. By minimizing data requirements while enhancing transparency in ML models, the CE feature provides a practical approach for developing efficient and accurate ML-based predictions tailored for small-data scenarios in materials science. This article is categorized under:
The cover image is based on the article Building Nucleosome Positioning Maps: Discovering Hidden Gems by Anna Panchenko et al., https://doi.org/10.1002/wcms.70029 .
Polypharmacy has become a routine practice in modern medicine, yet the risks of drug–drug interactions (DDIs) remain a critical challenge for patient safety. Given the vast number of possible drug combinations and the impracticality of exhaustive clinical testing, computational approaches have become indispensable for DDI prediction. Over the past 15 years, the field has shifted from handcrafted, similarity-based models to deep learning and graph neural networks (GNNs). Prediction tasks have also expanded from binary classification to multi-class, multi-label, cold-start, and higher-order settings. These reflect an emerging paradigm in both methodology and scope. Yet critical bottlenecks remain. Data sparsity, unreliable negatives, class imbalance, and source heterogeneity undermine robustness; models still struggle with generalization to unseen drugs, with mechanistic interpretability, and with capturing asymmetric or higher-order interactions. These limitations continue to impede translation into clinical and regulatory practice. In this Advanced Review, we critically assess methodological evolution, benchmark datasets, and emerging paradigms, including GNNs, large language models (including multimodal extensions), and generative AI, and examine their promises and limitations. We argue that next-generation progress hinges on unified multimodal and mechanism-aware frameworks, strategies for robust learning under cold-start and long-tail scenarios, and the integration of causal inference with generative approaches to enhance interpretability. By synthesizing past advances with forward-looking perspectives, this review outlines strategic pathways for accelerating the transition of DDI prediction toward intelligent, interpretable, and clinically actionable solutions. This article is categorized under:
The double-helical DNA of large eukaryotic genomes is tightly compacted within the tiny cell nucleus as a DNA–protein complex, chromatin. The universal elements of chromatin, nucleosome core particles (NCPs, 147 base pairs of DNA wrapped around an octamer of histone proteins), are connected by linker DNA of variable lengths into nucleosome arrays, which fold into various and dynamic higher-order structures. Since DNA is a highly negatively charged polyelectrolyte, electrostatic interactions of DNA with positively charged histones, other charged nuclear proteins, as well as with monovalent and multivalent cations, contributes decisively to the formation and folding of nucleosome arrays. The dimensions and timescales of cellular chromatin states and transformations necessitate a multiscale coarse-graining (CG) approach to understand their properties through computational modeling. In this review, we highlight the importance of electrostatics for NCP interactions and nucleosome fiber folding in vitro and in vivo, and argue that the inclusion of explicit ions is indispensable for accurate CG modeling of chromatin structure and dynamics. A summary of the existing CG mapping and force field setups is provided. A brief account of CG modeling studies in which salt dependency is approximated by the Debye–Hückel treatment is given. The primary focus is on the presentation of results from papers that include explicit monovalent and multivalent ionic species in CG simulations of nucleosomes and nucleosome arrays. Finally, we underline perspectives and challenges for future multiscale computational modeling of chromatin. This article is categorized under:
Principal component analysis (PCA) is a central tool for extracting essential information from complex datasets and has become widely used in the study of dynamical systems across disciplines. Its interdisciplinary relevance spans physics, chemistry, biology, computer science, and applied mathematics, where PCA and related approaches serve as gateways to understanding structure–function relationships, emergent behavior, and data-driven modeling. In the theoretical study of biomolecular systems using molecular dynamics (MD) simulations method, PCA filters high-dimensional trajectories into a reduced set of collective motions that elucidate conformational transitions and functional mechanisms. PCA provides an intuitive framework to connect statistical variance with dominant dynamical modes, a concept that extends naturally to the atomic scale of biomolecules. Modern developments integrate PCA with time-lagged methods, Markov state models, nonlinear dimensionality reduction, and machine learning techniques. These advances capture slow modes, rare events, and nonlinear manifolds, enriching the understanding of MD simulations results. A variety of computational packages now provide PCA-based analyses, supporting workflows from raw trajectory processing to visualization of free-energy landscapes and structural conformations. Applications range from probing peptide folding and protein domain motions to exploring collective dynamics in large assemblies. Since their first application more than 30 years ago to MD simulation, PCA-based methods continue to enhance the ability to analyze complex dynamical systems, offering a unifying statistical perspective that connects molecular simulations with interdisciplinary approaches to high-dimensional data analysis. This article is categorized under: