Artificial intelligence is transforming molecular and materials science, but its growing computational and data demands raise critical sustainability challenges. In this Perspective, we examine resource considerations across the AI-driven discovery pipeline–from quantum-mechanical (QM) data generation and model training to automated, self-driving research workflows–building on discussions from the “SusML workshop: Towards sustainable exploration of chemical spaces with machine learning” held in Dresden, Germany. In this context, the availability of large quantum datasets has enabled rigorous benchmarking and rapid methodological progress, while also incurring substantial energy and infrastructure costs. We highlight emerging strategies to enhance efficiency, including general-purpose machine learning (ML) models, multi-fidelity approaches, model distillation, and active learning. Moreover, incorporating physics-based constraints within hierarchical workflows, where fast ML surrogates are applied broadly and high-accuracy QM methods are used selectively, can further optimize resource use without compromising reliability. Equally important is bridging the gap between idealized computational predictions and real-world conditions by accounting for synthesizability and multi-objective design criteria, which is essential for practical impact. Finally, we argue that sustainable progress will rely on open data and models, reusable workflows, and domain-specific AI systems that maximize scientific value per unit of computation, enabling efficient and responsible discovery of technological materials and therapeutics.
Exact determination of the electronic density of molecules and materials would provide direct access to accurate bonded and non-bonded interatomic interactions by virtue of the Hellman-Feynman theorem. However, density-functional approximations (DFA) -- the workhorse methods for the electronic structure of atomistic systems -- only provide approximate and sometimes unreliable electron densities. Here we show that long-range van der Waals (vdW) dispersion interactions can visibly modify the charge density, scale nontrivially with system size, and in some cases cause polarization of charge density that exceeds that of the underlying semi-local density functional. We use the fully-coupled Many-Body Dispersion model to compute the vdW charge density by an appropriate real-space projection of the collective fluctuations of optimally-tuned coupled harmonic oscillators that constitute the model. Our analysis highlights the potential unreliability of post-hoc methods for vdW dispersion interactions, has implications for detecting interacting regions in large (bio)molecules, and enables the construction of more accurate DFAs and machine-learned force fields based on the electron density.
Recent advances in machinelearning force fields (MLFFs) have significantly extended the reach of atomistic simulations. Continuous progress in this field requires reliable reference data sets, accurate MLFF architectures, and efficient active learning strategies to enable robust modeling of complex molecular and material systems. Here we introduce aims-PAX, an expedited, multitrajectory active learning framework that streamlines the development of stable and accurate MLFFs. Designed for a wide range of researchers, aims-PAX offers a modular, high-performance workflow that couples diversified sampling with scalable training across CPU and GPU architectures. Integrated with the widely used ab initio code FHI-aims, the framework supports state-of-the-art ML models and data set generation using general-purpose (or "foundational") force-fields for rapid deployment in diverse systems. We demonstrate the capabilities of aims-PAX in various challenging tasks: creating data sets and models for highly flexible peptides, multiple organic molecules at once, explicitly solvated molecules, and for efficiently handling computationally demanding systems such as the CsPbI3 perovskite. We show that aims-PAX achieves a reduction of up to 3 orders of magnitude in the number of required reference calculations, automatically selects challenging systems within a given chemical space, facilitates simulation of solvated molecules with more than thousand atoms, while enabling a 10-fold speedup in active-learning time through optimized resource utilization. This positions aims-PAX as a powerful and versatile platform for next-generation atomistic simulations in both academic and industrial settings.
Interactions between objects can be classified as fundamental or emergent. Fundamental interactions are either extremely short-range or decay inversely with the separation distance, such as the Coulomb potential between charges or the gravitational attraction between masses. In contrast, emergent quantum van der Waals (vdW) and Casimir interactions decay considerably faster (R^-6 or R^-7) with distance R. Here we apply perturbative quantum electrodynamics (QED) to a many-body (MB) system of atoms modeled as charged harmonic oscillators, and reveal a persistent inverse-distance MB-QED interaction stemming from the coupling between virtual photons and molecular plasmons in the non-retarded regime. This interaction, scaling with the third power of the fine-structure constant, is reminiscent of the Lamb shift for a single atom. Although weaker than vdW forces, this MB-QED R^-1 interaction may substantially surpass gravitational attraction in future experiments probing quantum gravity at microscopic scales.
The Tkatchenko-Scheffler (TS) pairwise method to calculate dispersion interactions is a widely used approach to incorporate missing long-range van der Waals contributions in semilocal and hybrid density functional calculations. Despite numerous refinements of the approach to include many-body terms, the original formulation still remains highly relevant as an efficient and robust method, especially for organic and/or insulating materials. In 2018, Fedorov et al. reported updated van der Waals radii to the seminal work published in 2009. The present work examines the accuracy of the TS method with updated van der Waals radii (abbreviated as TS_2018), coupled with the semilocal Perdew-Burke-Ernzerhof density functional, for structural predictions of semiconducting and insulating materials in comparison to the non-local many-body dispersion method and the original TS method (TS_2009). Special attention is paid to materials containing alkali elements, for which the TS_2009 method exhibits a large overbinding, associated with potentially large errors in predicted atomic structures. We also consider a more narrow reformulation (TS_alkali) where only the the alkali atoms are corrected, so the method remains otherwise compatible with TS_2009. The binding energy curves of five alkali dimers are used to assess the TS_2009 and the TS_2018 methods in comparison to the random phase approximation. Using 45 inorganic solid compounds with available experimental reference data, as well as three widely studied, Cs-containing halide perovskites, CsPbX_3 (X = Cl, Br, I), we then examine the performance of the TS_2018 and TS_alkali approaches compared to TS_2009 and the beyond-pairwise, nonlocal many-body dispersion method; the latter found to give good results as well.
Classifying interactions is key in the physical sciences, and bonding mechanisms in matter-antimatter systems remain particularly enigmatic. Here we focus on a paradigmatic example of positronium hydride (PsH) dimer composed of two protons, two positrons, and four electrons, whose bonding nature has been previously described as either ionic, covalent, or van der Waals-like. Accurate quantum Monte Carlo calculations show that the two positrons occupy a delocalized molecular orbital that envelopes the two hydrogen anions and responds as a collective dipole to an applied electric field. This positronic bonding stems from quantum correlations that resemble a single covalent bond formed between negatively charged pseudo-nuclei, but with a bond strength commensurate with the traditional van der Waals interaction. Our findings suggest that the ability to form delocalized proto-bonds is a more general property of quantum systems, and could be present in a broader class of particles, antiparticles, and quasi-particles interacting with matter.
Realistic physical systems are characterised by emergent interactions across multiple length and time scales, posing a significant challenge for predictive machine learning (ML) models. Most scientific ML models focus on a narrow range of interactions. While machine learning force fields (MLFFs) offer near-quantum accuracy, the ubiquitous message-passing layers miss long-range many-body effects. Here we introduce the Multiscale Structural Ensemble (MuSE), a hierarchical model that uses Soft Coarse-Graining Pooling to construct coarse representations from smooth fractional assignments of atoms to coarse nodes, enabling MLFF modules to operate across multiple scales. MuSE is architecture-agnostic and coupled with SO3krates, MACE, and PaiNN MLFFs for both molecules and materials. We demonstrate the power of MuSE through Hessian-based benchmarks, folding trajectories for biomolecules, and energy profiles in molecule-graphene nanostructures, where MuSE accurately captures quantum-mechanical interactions at relevant scales – unlike other recent long-range ML models.
Machine learning (ML) approaches have drastically advanced the exploration of structure-property and property-property relationships in computer-aided drug discovery. A central challenge in this field is the identification of molecular descriptors that can effectively capture both geometric- and electronic structure-derived features, enabling the development of reliable and interpretable predictive models. While numerous descriptors focusing solely on structural characteristics have been recently proposed, improvements in model accuracy often come at the cost of increased computational demands, thereby restricting their practical applicability. To address this challenge, we introduce the "QUantum Electronic Descriptor" (QUED) framework, which integrates both structural and electronic data of molecules to develop ML regression models for property prediction. In doing so, a quantum-mechanical (QM) descriptor is derived from molecular and atomic properties computed using the semi-empirical density functional tight-binding (DFTB) method, which allows for efficient modelling of both small and large drug-like molecules. This descriptor is combined with inexpensive geometric descriptors-capturing two-body and three-body interatomic interactions-to form comprehensive molecular representations used to train Kernel Ridge Regression and XGBoost models. As a proof of concept, we validate QUED using the QM7-X dataset, which comprises equilibrium and non-equilibrium conformations of small drug-like molecules, demonstrating that incorporating electronic structure data notably enhances the accuracy of ML models for predicting physicochemical properties. For biological endpoints, we find that QM properties provide some predictive value for toxicity and lipophilicity prediction, as assessed using the TDCommons-LD50 and the MoleculeNet benchmark datasets. Moreover, a SHapley Additive exPlanations (SHAP) analysis of the toxicity and lipophilicity predictive models reveals that molecular orbital energies and DFTB energy components are among the most influential electronic features. Hence, our work underscores the importance of incorporating QM descriptors to enhance both the accuracy and interpretability of ML models for predicting multiple properties relevant to pharmaceutical and biological applications.
Recent advances in machine learning force fields (MLFFs) are revolutionizing molecular simulations by bridging the gap between quantum-mechanical (QM) accuracy and the computational efficiency of mechanistic potentials. However, the development of reliable MLFFs for biomolecular systems remains constrained by the scarcity of high-quality, chemically diverse QM datasets that span all of the major classes of biomolecules expressed in living cells. Crucially, such a comprehensive dataset must be computed using non-empirical or minimally empirical approximations to solving the Schrödinger equation. To address these limitations, we introduce the QCell dataset – a curated collection of 525k new QM calculations for biomolecular fragments encompassing carbohydrates, nucleic acids, lipids, dimers, and ion clusters. QCell complements existing datasets, bringing the total number of available data points to 41 million molecular systems, all calculated using hybrid density functional theory with nonlocal many-body dispersion interactions, as captured by the PBE0+MBD(-NL) level of quantum mechanics. The QCell dataset therefore provides a valuable resource for training next-generation MLFFs capable of modeling the intricate interactions that govern biomolecular dynamics beyond small molecules and proteins.
This study examines the effect of combining quantum mechanical (QM) descriptors with traditional structural descriptors in predicting ADMET (absorption, distribution, metabolism, excretion, and toxicity) properties for molecular drug candidates. We focus on four end points: HLM (human liver microsomes) stability, permeability, solubility, and hERG (human ether-a-go-go-related gene) inhibition. First, we studied the prediction of molecules in the same chemical space as the training data. Our results show that incorporating QM descriptors enhances prediction accuracy across all end points. We find that the most significant improvement in permeability is achieved, with the QM dispersion energy proving to be the second most important feature. In general, feature importance analysis revealed that QM descriptors are ranked among the top five features, providing more comprehensive insights than structural descriptors alone. Next, we studied predicting molecules in chemical spaces distinct from the training data. Here, QM descriptors exhibited superior extrapolation capabilities, particularly for permeability and hERG data sets. In an extrapolation scenario where the training data set size was systematically reduced, QM descriptors maintained higher stability than structural descriptors in prediction. A hybrid approach combining all QM descriptors with key structural descriptors improved predictive accuracy, reduced variation in results, and yielded greater robustness in small data sets. This work highlights the value of QM descriptors for ADMET predictions, especially in challenging extrapolation tasks and scenarios with limited training data.
Computational blind challenges offer critical, unbiased opportunities to assess and accelerate scientific progress, as demonstrated by a breadth of breakthroughs over the past decade. We report the outcomes and key insights from an open science community blind challenge focused on computational methods in drug discovery, using lead optimization data from the AI-driven Structure-enabled Antiviral Platform Discovery Consortium's pan-coronavirus antiviral discovery program, in partnership with Polaris and the OpenADMET project. This collaborative initiative invited global participants from both academia and industry to develop and apply computational methods to predict the biochemical potency and crystallographic ligand poses of small molecules against key coronavirus targets, Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2) and Middle East Respiratory Syndrome Coronavirus (MERS-CoV) main protease (Mpro), as well as multiple ADMET assay end points, using previously undisclosed comprehensive experimental drug discovery data sets as benchmarks. By evaluating submissions across multiple tasks and compounds, we established performance leaderboards and conducted meta-analyses to assess methodological strengths, common pitfalls, and areas for improvement. This analysis provides a foundation for best practices in real-world machine learning evaluation, grounded in community-driven benchmarking. We also highlight how next-generation platforms, such as Polaris, enable rigorous challenge design, embedded evaluation frameworks, and broad community engagement. This paper reports the collective findings of the challenge, offering a high-level overview of the data, evaluation infrastructure, and top-performing strategies. We further provide context and support for the accompanying papers authored by the challenge participants in this special issue, which explore individual approaches in greater depth. Together, these contributions aim to advance reproducible, trustworthy, and high-impact computational methods in drug discovery, and to explore best practices and pitfalls in future blind challenge design and execution, including planned initiatives for the OpenADMET project.
Van der Waals (vdW) forces and the associated Hamaker constants underpin a broad spectrum of phenomena in biology, chemistry, materials science, and nanotechnology, shaping processes from protein folding to the stability of two-dimensional (2D) materials. Achieving precise nanoscale quantification of these interactions remains a challenge, as conventional approaches lack the spatial resolution to resolve localized vdW effects in complex systems. Atomic force microscopy (AFM), with its nanometer sharp probe, has overcome many of these limitations, enabling high-resolution and nondestructive measurement of vdW forces and the Hamaker constants. Here, the physics of vdW interactions and recent advances in AFM methodologies that have transformed their quantification is reviewed. Applications span 2D materials, where AFM reveals layer-layer and layer-substrate interactions critical for device performance, and biological systems, where vdW forces govern protein assembly, viral mutation, and cellular transformations. Emerging integration of AFM with molecular and theoretical frameworks for intermolecular interactions based on quantum mechanics further expand the frontier, offering unprecedented accuracy and insights into nanoscale phenomena. Perspectives and prospects of AFM-based vdW force detection is discussed, highlighting its potential in providing a high-resolution and noninvasive quantitative instrument for the development of novel applications in materials science, molecular biology, and quantum technologies.
We present a structured force reformulation of the many-body dispersion (MBD) model that enables a physically consistent decomposition of forces into pairwise components. By introducing a many-body correlation matrix that scales dipole–dipole interactions, we derive unified expressions for the MBD energy, force, and Hessian. This reformulation reveals a natural structure for effective atom–atom force decomposition and provides a promising foundation for interpretable analysis and machine learning surrogate modeling of MBD interactions.
The ΛCDM model successfully explains a wide range of cosmological observations; however, persistent discrepancies most notably the H_0 tension between early and late time measurements challenge its completeness. No proposed extension has yet resolved this tension while retaining the overall success of ΛCDM. In this work, we investigate whether the H_0 tension can be associated with a specific epoch in the cosmic expansion history and identify the redshift range most relevant for understanding its origin. In addition to the cosmological constant, we consider three phenomenological models based on general parametrizations of key quantities governing cosmic expansion: the dark energy (DE) equation of state, the DE pressure density, and the scale factor. Using early time Planck data and late time Pantheon+ (with and without SH0ES calibration) and DESI measurements, we constrain model parameters and examine the evolution of the Hubble parameter H(z). We find that ΛCDM exhibits discrepancies across all redshifts, whereas the other models shift the dominant deviations toward low redshifts. Among the models considered, the pressure density parametrization alleviates the H_0 tension, reducing it to ∼ 2.7σ, while the other models do not provide significant improvement. A detailed analysis of DESI DR2 data further reveals notable deviations in H(z) at z=0.51 and 0.706, whereas higher redshift measurements remain consistent within 1σ. These results suggest that late-time modifications primarily reshape the redshift dependence of the mismatch in H(z) rather than fully resolve it, in the absence of systematic effects. Furthermore, the reconstructed DE dynamics exhibit qualitatively distinct behaviors across parametrizations, highlighting a persistent inconsistency between early and late Universe probes in describing the nature of DE.
Biomolecular thermodynamics and spectroscopy depend on relative conformer energies, local curvatures, and collective dipole fluctuations on the potential-energy surface. Conventional molecular mechanics force fields enable large-scale simulations, but their fixed functional forms can misrepresent infrared intensities, mode character, and environment-dependent vibrational response. Here we assess general-purpose machine-learned force fields across small molecules, finite-temperature infrared spectra, gas-phase peptides, and monomeric, oligomeric, and solvated protein assemblies. To enable this analysis, we introduce QVib, a dataset of 293 molecules and 1365 conformers, together with peptide amide-band benchmarks and p53 oligomerization-domain models, to evaluate vibrational transferability from DFT references to experimental spectra. Across these systems, machine-learned force fields substantially improve over molecular mechanics in reproducing DFT-level forces, vibrational frequencies, densities of states, mode eigenvectors, conformational energetics, and experimental infrared spectra. Among models with explicit long-range electrostatics, SO3LR provides the most favourable accuracy-cost balance for the biomolecular systems considered. These results show that machine-learned force-field dynamics can recover collective, environment-dependent vibrational landscapes at near-DFT fidelity, enabling spectroscopically validated biomolecular simulations at force-field-like cost.
Van der Waals (vdW) interactions are essential for describing molecules and materials, from drug design and catalysis to battery applications. These omnipresent interactions must also be accurately included in machine-learned force fields. The many-body dispersion (MBD) method stands out as one of the most accurate and transferable approaches to capture vdW interactions, requiring only atomic $C_6$ coefficients and polarizabilities as input. We present MBD-ML, a pretrained message passing neural network that predicts these atomic properties directly from atomic structures. Through seamless integration with libMBD, our method enables the immediate calculation of MBD-inclusive total energies, forces, and stress tensors. By eliminating the need for intermediate electronic structure calculations, MBD-ML offers a practical and streamlined tool that simplifies the incorporation of state-of-the-art vdW interactions into any electronic structure code, as well as empirical and machine-learned force fields.
Accurate prediction of many-body dispersion (MBD) interactions is essential for understanding the van der Waals forces that govern the behavior of many complex molecular systems. However, the high computational cost of MBD calculations limits their direct application in large-scale simulations. In this work, we introduce a machine learning surrogate model specifically designed to predict MBD forces in polymer melts, a system that demands accurate MBD description and offers structural advantages for machine learning approaches. Our model is based on a trimmed SchNet architecture that selectively retains the most relevant atomic connections and incorporates trainable radial basis functions for geometric encoding. We validate our surrogate model on datasets from polyethylene, polypropylene, and polyvinyl chloride melts, demonstrating high predictive accuracy and robust generalization across diverse polymer systems. In addition, the model captures key physical features, such as the characteristic decay behavior of MBD interactions, providing valuable insights for optimizing cutoff strategies. Characterized by high computational efficiency, our surrogate model enables practical incorporation of MBD effects into large-scale molecular simulations.
Accurately modeling noncovalent interactions (NCIs) involving charged systems remains an outstanding challenge in density functional theory (DFT), with implications across natural and life sciences, engineering, e.g., in biochemistry, catalysis, and materials science. For these interactions, the interplay between electrostatics, polarization, and dispersion leads to systematic errors of up to tens of kilocalories per mole in standard dispersion-enhanced DFT methods. We solve this problem by introducing (r2SCAN+MBD)@HF, a DFT method without empirically fitted parameters that combines the r2SCAN functional and many-body dispersion, both evaluated on Hartree-Fock densities. We show that the unique synergy of these three components enables balanced treatment of short- and long-range correlation, which is crucial for accurate description of NCIs involving charged systems. Evaluations on standard benchmarks show that (r2SCAN+MBD)@HF significantly improves accuracy for NCIs involving charged systems while maintaining robust performance for neutral systems. Further tests on the Metal Ion Protein Clusters dataset introduced here demonstrate its improved accuracy for metal-protein interactions, including cases where the metal ion is surrounded by negatively charged ligands for which standard DFT can completely break down. Given the ubiquity of such interactions, (r2SCAN+MBD)@HF is broadly applicable from biochemistry and materials science, including for generating high-quality data to train machine-learning force fields.
We present the first open access version of the QMeCha (Quantum MeCha /'mεkə/) code, a quantum Monte Carlo package developed to study many-body interactions between different types of quantum particles, with a modular and easy-to-expand structure. QMeCha is now available under a CC BY-NC-ND license through the repository github.com/QMeCha. The present code has been developed to solve the Hamiltonian of a system that can include nuclei and fermions of different mass and charge, e.g., electrons and positrons, embedded in an environment of classical charges and quantum Drude oscillators. To approximate the ground state of this many-particle operator, the code features different wavefunctions. For the fermionic particles, beyond the traditional Slater determinant, QMeCha also includes geminal functions, such as the Pfaffian, and presents different types of explicit correlation terms in Jastrow factors. The classical point charges and quantum Drude oscillators, described through different variational Ansätze, are used to model a molecular environment capable of explicitly describing dispersion, polarization, and electrostatic effects experienced by the nuclear and fermionic subsystems. To integrate these wavefunctions, efficient variational Monte Carlo and diffusion Monte Carlo protocols have been developed, together with a robust wavefunction optimization procedure that features correlated sampling.
Isolated atoms as well as molecules at equilibrium are presumed to be simple from the point of view of quantum computational complexity. Here, we show that the process of chemical bond formation is accompanied by a marked increase in the quantum complexity of the electronic ground state. By studying the hydrogen dimer H_{2} as a prototypical example, we demonstrate that when two hydrogen atoms form a bond, a specific measure of quantum complexity exhibits a pronounced peak that closely follows the behavior of the binding energy. This measure of quantum complexity, known as magic in the quantum information literature, reflects how difficult it is to simulate the state using classical methods. We show that the observations for H_{2} also hold for a collection of other dimers, including the weakly bonded diatomic helium dimer He_{2}. This observation suggests that regions of strong bonding formation or breaking are also regions of enhanced intrinsic quantum complexity. This insight suggests a connection of quantum information measures to chemical reactivity and advocates the use of stretched molecules as a quantum computational resource.