Mixtures of chemical ingredients, such as formulations, are ubiquitous in materials science, but optimizing their properties remains challenging due to the vast design space. Computational approaches offer a promising solution to traverse this space while minimizing trial-and-error experimentation. Using high-throughput classical molecular dynamics simulations, we generated a comprehensive dataset of over 30,000 solvent mixtures to evaluate three machine learning approaches that connect molecular structure and composition to property: formulation descriptor aggregation (FDA), formulation graph (FG), and Set2Set-based method (FDS2S). Our results demonstrate that our new FDS2S approach outperforms other approaches in predicting simulation-derived properties. Formulation-property relationships can reveal important substructures and identify promising formulations at least two to three times faster than random guessing. The models show robust transferability to experimental datasets, accurately predicting properties across energy, pharmaceutical, and petroleum applications. Our research demonstrates the utility of high-throughput simulations and machine learning tools to design formulations with promising properties.
In materials science, accurately computing properties like viscosity, melting point, and glass transition temperatures solely through physics-based models is challenging. Data-driven machine learning (ML) also poses challenges in constructing ML models, especially in the material science domain where data is limited. To address this, we integrate physics-informed descriptors from molecular dynamics (MD) simulations to enhance the accuracy and interpretability of ML models. Our current study focuses on accurately predicting viscosity in liquid systems using MD descriptors. In this work, we curated a comprehensive dataset of over 4,000 small organic molecules’ viscosities from scientific literature, publications, and online databases. This dataset enabled us to develop quantitative structure–property relationships (QSPR) consisting of descriptor-based and graph neural network models to predict temperature-dependent viscosities for a wide range of viscosities with considerable accuracy. The QSPR models reveal that including MD descriptors improves prediction accuracies of experimental viscosities, particularly at the small data set scale of fewer than a thousand data points. Furthermore, feature importance tools reveal that intermolecular interactions captured by MD descriptors are most important for accurate viscosity predictions. Finally, the QSPR models can accurately capture the inverse relationship between viscosity and temperature for six battery-relevant solvents, some of which were not included in the original data set. Our research highlights the effectiveness of incorporating MD descriptors into QSPR models, which leads to improved accuracy for properties that are difficult to predict when using physics-based models alone or when limited data is available.
Epik version 7 is a software program that uses machine learning for predicting the pKa values and protonation state distribution of complex, drug-like molecules. Using an ensemble of atomic graph convolutional neural networks (GCNNs) trained on over 42,000 pKa values across broad chemical space from both experimental and computed origins, the model predicts pKa values with 0.42 and 0.72 log unit median absolute and RMS errors, respectively, across seven test sets. Epik version 7 also generates protonation states and recovers 95% of the most populated protonation states compared to previous versions. Requiring on average only 47 ms per ligand, Epik version 7 is rapid and accurate enough to evaluate protonation states for crucial molecules and prepare ultra-large libraries of compounds to explore vast regions of chemical space. The simplicity of and time required for the training allows for the generation of highly accurate models customized to a program’s specific chemistry.
The blood-brain barrier (BBB) plays a critical role in preventing harmful endogenous and exogenous substances from penetrating the brain. Optimal brain penetration of small-molecule central nervous system (CNS) drugs is characterized by a high unbound brain/plasma ratio (Kp,uu). While various medicinal chemistry strategies and in silico models have been reported to improve BBB penetration, they have limited application in predicting Kp,uu directly. We describe a physics-based computational approach, a quantum mechanics (QM)-based energy of solvation (E-sol), to predict Kp,uu. Prospective application of this method in internal CNS drug discovery programs highlights the utility and accuracy of this new method, which showed a categorical accuracy of 79% and an R2 of 0.61 from a linear regression model.