Accurately measuring compound binding affinities is key to driving the pharmaceutical development process. Rigorous physics-based in silico approaches, particularly alchemical free energy methods, have become a gold-standard tool for estimating compound affinity changes. Here we present the results of a large-scale precompetitive collaborative assessment of relative binding free energy (RBFE) calculations generated by 15 pharmaceutical companies. We evaluate an open-source and MIT-licensed RBFE protocol from the Open Free Energy (OpenFE) ecosystem across both public and blinded private datasets, encompassing over 1,700 ligands in total. For the public dataset, the weighted RMSE across the 58 systems was 1.73(1.53)(1.96) kcal/mol, with 10 of the systems reaching sub-kcal/mol accuracy. For the private dataset, the weighted RMSE across the 37 systems was 2.44(1.94)(3.06) kcal/mol, with only 2 of the systems reaching sub-kcal/mol accuracy, reflecting the increased complexity of real-world drug discovery. The protocol's performance was system-dependent, with no single dominant error source, indicating that accuracy is primarily influenced by input quality and transformation type. Overall, these benchmark results are encouraging and indicate that OpenFE is ready for large-scale industrial applications, with an "out-of-the-box" accuracy that approaches that of commercial solutions. While comparison against published FEP+ results, which were obtained after manual parameter optimization, shows comparable ranking statistics, a gap remains in error statistics, for which we outline possible paths toward improvement. The protocol meets key criteria required for production use in an industrial setting: it shows robust performance, generates reproducible results, and achieves both sufficient throughput and rapid convergence.
A wide range of density functional methods and basis sets are available to derive the electronic structure and properties of molecules. Quantum mechanical calculations are too computationally intensive for routine simulation of molecules in the condensed phase, prompting the development of computationally efficient force fields based on quantum mechanical data. Parametrizing general force fields, which cover a vast chemical space, necessitates generating sizable quantum mechanical datasets with optimized geometries and torsion scans. To achieve this efficiently, it is crucial to choose a quantum mechanical method that balances computational cost and accuracy. In this study we seek to assess the accuracy of quantum mechanical theory for specific properties such as conformer energies, electrostatic properties, and torsion energetics. To comprehensively evaluate various methods, we focus on a representative set of 59 diverse small molecules, comparing approximately 24 combinations of functional and basis sets against the reference level coupled cluster calculations at complete basis set limit.
alchemlyb is an open-source Python software package for the analysis of alchemical free energy calculations, an important method in computational chemistry and biology, most notably in the field of drug discovery (Merz et al., 2010). Its functionality contains individual composable building blocks for all aspects of a full typical free energy analysis workflow, starting with the extraction of raw data from the output of diverse molecular simulation packages, moving on to data preprocessing tasks such as decorrelation of time series, using various estimators to derive free energy estimates from simulation samples, and finally providing quality analysis tools for data convergence checking and visualization. alchemlyb also contains high-level end-to-end workflows that combine multiple building blocks into a user-friendly analysis pipeline from the initial data input stage to the final results. This workflow functionality enhances accessibility by enabling researchers from diverse scientific backgrounds, and not solely computational chemistry specialists, to use alchemlyb effectively.
A simple space elevator consists of a single tether extending well beyond geosynchronous altitude and a payload-carrying device which grips and climbs the tether. A friction-based, opposing wheel climber was judged most likely to be constructed with present-day technology and it appears that mass-production of the tether material is also within reach. The physical conditions at the interface between the climber wheels and tether determine first of all the possibility of climbing and then the design parameters of the tether. Conditions such as lifting torque, tensile, compressive and shear strength, friction, interface temperature, thermal conductivity and radiative cooling were examined and used to set minimum requirements for the tether material. Graphene superlaminate (GSL), consisting of layers of single crystal graphene, appears to be an excellent tether material with a sufficiently high tensile strength. An increase in its inter-layer cross-bonding and a larger mutual coefficient of friction with the climber wheel material would allow it to satisfy the climbing conditions. A final determination of the suitability of GSL requires the measurement of a number of, as yet unknown, material properties. A list of such measurements is proposed and a partial list of trade studies and iterations of design for the tether are provided.
We introduce the Open Force Field (OpenFF)~2.0.0 small molecule force field for drug-like molecules, code-named Sage, which builds upon our previous iteration, Parsley. OpenFF force fields are based on direct chemical perception, which generalizes easily to highly diverse sets of chemistries based on substructure queries. Like the previous OpenFF iterations, the Sage generation of OpenFF force fields was validated in protein-ligand simulations to be compatible with AMBER biopolymer force fields. In this paper we detail the methodology used to develop this force field, as well as the innovations and improvements introduced since the release of Parsley 1.0.0. One particularly significant feature of Sage is a set of improved Lennard-Jones (LJ) parameters retrained against condensed phase mixture data, the first refit of LJ parameters in the OpenFF small molecule force field line. Sage also includes valence parameters refit to a larger database of quantum chemical calculations than previous versions, as well as improvements in how this fitting is performed. Force field benchmarks show improvements in general metrics of performance against quantum chemistry reference data such as root mean square deviations (RMSD) of optimized conformer geometries, torsion fingerprint deviations (TFD), and improved relative conformer energetics (ΔΔ𝐸). We present a variety of benchmarks for these metrics against our previous force fields as well as in some cases other small molecule biomolecular force fields. Sage also demonstrates improved performance in estimating physical properties, including comparison against experimental data from various thermodynamic databases for small molecule properties such as Δ𝐻_𝑚𝑖𝑥, ρ(𝑥), Δ𝐺_𝑠𝑜𝑙𝑣 and Δ𝐺_𝑡𝑟𝑎𝑛𝑠. Additionally, we benchmarked against protein-ligand binding free energies (Δ𝐺_𝑏𝑖𝑛𝑑), where Sage yields results statistically similar to previous force fields. All the data is made publicly available along with complete details on how to reproduce the training results at https://github.com/openforcefield/openff-sage.
Software to more rapidly and accurately predict protein--ligand binding affinities is of high interest for early-stage drug discovery, and physics-based methods are among the most widely used technologies for this purpose. The accuracy of these methods depends critically on the accuracy of the potential functions they use. Potential functions are typically trained against a combination of quantum chemical and experimental data. However, although binding affinities are among the most important quantities to predict, experimental binding affinities have not to date been integrated into the experimental dataset used to train potential functions. In recent years, the use of host--guest complexes as simple and tractable models of binding thermodynamics has gained popularity due to their small size and simplicity, relative to protein--ligand systems. Host--guest complexes can also avoid ambiguities that arise in protein--ligand systems, such as uncertain protonation states. Thus, experimental host--guest binding data are an appealing additional data type to integrate into the experimental dataset used to optimize potential functions. Here, we report the extension of the Open Force Field Evaluator framework to enable the systematic calculation of host--guest binding free energies and their gradients with respect to force field parameters, coupled with the curation of 126 host--guest complexes with available experimental binding free energies. As an initial application of this novel infrastructure, we optimized generalized Born (GB) cavity radii for the OBC2 GB implicit solvent model against experimental data for 36 host--guest systems. This refitting led to a dramatic improvement in accuracy for both the training set and a separate test set with 90 additional host--guest systems. The optimized radii also showed encouraging transferability from host--guest systems to 59 protein-ligand systems. However, the new radii are significantly smaller than the baseline radii and lead to excessively favorable hydration free energies (HFE). Thus, users of the OBC2 GB model currently may choose between GB cavity radii that yield more accurate binding affinities or GB cavity radii that yield more accurate HFEs. We suspect that achieving good accuracy on both will require more far-reaching adjustments to the GB model. We note that binding free energy calculations using the OBC2 model in OpenMM gain about a 10x speedup relative to corresponding explicit solvent calculations, suggesting a future role for implicit solvent absolute binding free energy (ABFE) calculations in virtual compound screening. This study proves the principle of using host--guest systems to train potential functions that are transferrable to protein--ligand systems, and provides an infrastructure that enables a range of applications.
Proteins containing non-canonical residues play important roles in cellular biology and drug design. Developing force field (FF) parameters for these proteins is challenging and time-consuming, and often no straightforward and automated process can extend an existing FF to model new chemistries. Additionally, parameters transferred from general FFs to biomolecule FFs are often trained differently than biomolecule parameters—e.g., torsions trained to different levels of quantum chemical (QC) theory—creating inconsistencies between the treatment of canonical and non-canonical residues.
Machine learning potentials are an important tool for molecular simulation, but their development is held back by a shortage of high quality datasets to train them on. We describe the SPICE dataset, a new quantum chemistry dataset for training potentials relevant to simulating drug-like small molecules interacting with proteins. It contains over 1.1 million conformations for a diverse set of small molecules, dimers, dipeptides, and solvated amino acids. It includes 15 elements, charged and uncharged molecules, and a wide range of covalent and non-covalent interactions. It provides both forces and energies calculated at the ωB97M-D3(BJ)/def2-TZVPPD level of theory, along with other useful quantities such as multipole moments and bond orders. We train a set of machine learning potentials on it and demonstrate that they can achieve chemical accuracy across a broad region of chemical space. It can serve as a valuable resource for the creation of transferable, ready to use potential functions for use in molecular simulations.
The conditions at the interface between the space elevator tether and its climber determine the requirements of the tether material and the climber design. Graphene super-laminate was found to have sufficient tensile strength to support the tether and climber, but improvements of the tether manufacturing process are required in order to increase friction and shear strength. Based on these requirements, a 20 metric ton climber was designed, could carry a 9 ton payload and could be built in the near-future.
Accurate small molecule force fields are crucial for predicting thermodynamic and kinetic properties of drug-like molecules in biomolecular systems. Torsion parameters, in particular, are essential for determining conformational distribution of molecules. However, they are usually fit to computationally expensive quantum chemical torsion scans and generalize poorly to different chemical environments. Torsion parameters should ideally capture local through-space non-bonded interactions such as 1-4 steric and electrostatics and non-local through-bond effects such as conjugation and hyperconjugation. Non-local through-bond effects are sensitive to remote substituents and are a contributing factor to torsion parameters poor transferability. Here we show that fractional bond orders such as the Wiberg Bond Order (WBO) are sensitive to remote substituents and correctly captures extent of conjugation and hyperconjugation. We show that the relationship between WBO and torsion barrier heights are linear and can therefore serve as a surrogate to QC torsion barriers, and to interpolate torsion force constants. Using this approach we can reduce the number of computationally expensive QC torsion scans needed while maintaining accurate torsion parameters. We demonstrate this approach to a set of substituted benzene rings.
Force fields form the basis for classical molecular simulations and their accuracy is crucial for the quality of, for instance, protein-ligand binding simulations in drug discovery. The huge diversity of small molecule chemistry makes it a challenge to build and parameterize a suitable force field. The Open Force Field Initiative is a combined industry and academic consortium developing a state-of-the-art small molecule force field. In this report industry members of the consortium worked together to objectively evaluate the performance of the force fields (referred to here as OpenFF) produced by the initiative on a combined public and proprietary dataset of 19,653 relevant molecules selected from their internal research and compound collections. This evaluation was important because it was completely blind; at most partners, none of the molecules or data were used in force field development or testing prior to this work. We compare the Open Force Field "Sage" version 2.0.0 and "Parsley" version 1.3.0 with GAFF-2.11-AM1BCC, OPLS4 and SMIRNOFF99Frosst. We analyzed force field-optimized geometries and conformer energies compared to reference quantum mechanical data. We show that OPLS4 performs best, and the latest Open Force Field release shows a clear improvement compared to its predecessors. The performance of established force fields such as GAFF-2.11 was generally worse. While OpenFF researchers were involved in building the benchmarking infrastructure used in this work, benchmarking was done entirely in-house within industrial organizations and the resulting assessment is reported here. This work assesses the force field performance using separate benchmarking steps, external datasets, and involving external research groups. This effort may also be unique in terms of the number of different industrial partners involved, with 10 different companies participating in the benchmark efforts.
The development of accurate transferable force fields is key to realizing the full potential of atomistic modeling in the study of biological processes such as protein-ligand binding for drug discovery. State-of-the-art transferable force fields, such as those produced by the Open Force Field Initiative, use modern software engineering and automation techniques to yield accuracy improvements. However, force field torsion parameters, which must account for many stereoelectronic and steric effects, are considered to be less transferable than other force field parameters and are therefore often targets for bespoke parametrization. Here, we present the Open Force Field QCSubmit and BespokeFit software packages that, when combined, facilitate the fitting of torsion parameters to quantum mechanical reference data at scale. We demonstrate the use of QCSubmit for simplifying the process of creating and archiving large numbers of quantum chemical calculations, by generating a dataset of 671 torsion scans for druglike fragments. We use BespokeFit to derive individual torsion parameters for each of these molecules, thereby reducing the root-mean-square error in the potential energy surface from 1.1 kcal/mol, using the original transferable force field, to 0.4 kcal/mol using the bespoke version. Furthermore, we employ the bespoke force fields to compute the relative binding free energies of a congeneric series of inhibitors of the TYK2 protein, and demonstrate further improvements in accuracy, compared to the base force field (MUE reduced from 0.560.390.77 to 0.420.280.59 kcal/mol and R2 correlation improved from 0.720.350.87 to 0.930.840.97).
Global climate action is the grand challenge of the 21st century. Large reductions in greenhouse gas emissions are needed, with net-zero CO2 emissions by mid-century a requirement for many scenarios that restrain global warming below 1.5 degrees C by 2100. This will require massive deployment of renewables such as ground-based solar and wind, increased long-distance grid transmission capacity, and electrification of much of the global economy. Increased electrification, coupled with rising standards of living in developing countries, will drive ever-larger electricity demand to 2050 and beyond. And the intermittency, land-use, and transmission requirements of ground-based solar and wind present challenges to large-scale deployment to meet Net Zero targets. Space-based solar power (SSP) presents an enticing alternative, and with full reusability of upcoming heavy launch vehicles and modular SSP architectures, the economics of deployment are expected to become favorable in the coming decades. Efforts exist today to deploy demonstrators this decade, with initial production systems by 2050. However, deploying SSP at the scale required to address rising, global electricity demand will itself require orders-of-magnitude greater launch frequencies. If heavy launch frequency does not increase quickly enough, due to environmental concerns, launch site competition, or other bottlenecks, the dropping cost of launch will not be sufficient to enable SSP for global impact this century. A space elevator (SE) development program like that proposed by the International Space Elevator Consortium (ISEC) offers an alternative delivery mechanism for scaling SSP beyond 2050, hedging the risk of reliance on a single delivery system while reducing the time required for large-scale SSP deployment. To de-risk possible futures in which SSP is a major part of the global energy mix this century, a dual space access architecture is proposed, leveraging the strengths of heavy launch and SEs together. Such an architecture would secure the future for green energy on Earth, in 2050 and beyond.
The Open Force Field (OpenFF) Initiative develops open, reproducible force fields for atomistic molecular simulations as well as automated infrastructure and systematic methodology for their parametrization and validation. Here we describe the extension of the OpenFF infrastructure to support biopolymers, enabling simulations of both small molecules and biopolymers with a self-consistent force field. The OpenFF workflow includes OpenFF QCSubmit for automated submission and curation of quantum chemical (QC) datasets via the Molecular Sciences Software Institute QCArchive project; OpenFF Toolkit for preparation of molecular topologies and parameter assignment using direct chemical perception via SMIRKS Native Open Force Field (SMIRNOFF) typing; OpenFF Evaluator for reproducible estimation of condensed phase physical properties from simulations; and ForceBalance for robust parameter optimization targeting both QC and physical property data.
SPICE (Small-Molecule/Protein Interaction Chemical Energies) is a collection of quantum mechanical data for training potential functions. The emphasis is particularly on simulating drug-like small molecules interacting with proteins. It is described in this publication: Peter Eastman, Pavan Kumar Behara, David L. Dotson, Raimondas Galvelis, John E. Herr, Josh T. Horton, Yuezhi Mao, John D. Chodera, Benjamin P. Pritchard, Yuanqing Wang, Gianni De Fabritiis, and Thomas E. Markland. "SPICE, A Dataset of Drug-like Molecules and Peptides for Training Machine Learning Potentials." https://doi.org/10.48550/arXiv.2209.10702 (2022). The HDF5 file is structured as follows. There is one top level group for each unique molecule or cluster. The name of each group is either a PubChem Substance ID (for PubChem molecules), an amino acid sequence (for dipeptides and solvated amino acids), or a SMILES string (for everything else). Each group contains the following datasets. N is the number of atoms in the molecule and M is the number of conformations. (Some groups may be missing some of them, for example if MBIS failed to converge.) subset: The name of the data subset the molecule is from. smiles: The canonical SMILES string for the molecule. It includes explicit hydrogens and atom indices. atomic_numbers: Array of length N containing the atomic number of every atom. They are ordered following the indices in the SMILES string. conformations: Array of shape (M, N, 3) containing the atomic coordinates for every conformation. formation_energy: Array of length M containing the total energy of each conformation, minus the reference energies of the individual atoms when infinitely separated. This is the most useful energy for most purposes, since it contains all energy components that vary with atom positions but removes the large constant part corresponding to the internal energies of individual atoms. dft_total_energy: Array of length M containing the energy of each conformation. dft_total_gradient: Array of shape (M, N, 3) containing the gradient of the energy with respect to the atomic coordinates. mbis_charges: Array of shape (M, N, 1) containing the MBIS charge of each atom. mbis_dipoles: Array of shape (M, N, 3) containing the MBIS dipole of each atom. mbis_quadrupoles: Array of shape (M, N, 3, 3) containing the MBIS quadrupole of each atom. mbis_octupoles: Array of shape (M, N, 3, 3, 3) containing the MBIS octupole of each atom. scf_dipoles: Array of shape (M, 3) containing the dipole of each molecule. scf_quadrupole: Array of shape (M, 3, 3) containing the quadrupole of each molecule. mayer_indices: Array of shape (M, N, N) containing the Mayer bond indices. wiberg_lowdin_indices: Array of shape (M, N, N) containing the Wiberg bond indices using orthogonal Löwdin orbitals. All values are in atomic units. Distances are in bohr and energies in hartree.
Community efforts in the computational molecular sciences (CMS) are evolving toward modular, open, and interoperable interfaces that work with existing community codes to provide more functionality and composability than could be achieved with a single program. The Quantum Chemistry Common Driver and Databases (QCDB) project provides such capability through an application programming interface (API) that facilitates interoperability across multiple quantum chemistry software packages. In tandem with the Molecular Sciences Software Institute and their Quantum Chemistry Archive ecosystem, the unique functionalities of several CMS programs are integrated, including CFOUR, GAMESS, NWChem, OpenMM, Psi4, Qcore, TeraChem, and Turbomole, to provide common computational functions, i.e., energy, gradient, and Hessian computations as well as molecular properties such as atomic charges and vibrational frequency analysis. Both standard users and power users benefit from adopting these APIs as they lower the language barrier of input styles and enable a standard layout of variables and data. These designs allow end-to-end interoperable programming of complex computations and provide best practices options by default.