We present a hybrid semiempirical density functional tight-binding (DFTB) model with a machine learning neural network potential as a correction to the repulsive term. This hybrid model, termed machine learning tight-binding (MLTB), employs the standard self-consistent charge (SCC) DFTB formalism as a baseline, enhanced by the HIP-NN potential as an effective many-body correction for short-range pairwise repulsive interactions. The MLTB model demonstrates significantly improved transferability and extensibility compared to the SCC-DFTB and HIP-NN models. This work provides a practical computational framework for developing reliable SCC-DFTB models with additional many-body corrections that more closely approach the DFT level of accuracy. We illustrate this method with the development of an accurate model for the thorium-oxygen system, applied to the study of its nanocluster structures (ThO2)n.
Advances in machine learning have given rise to a plurality of data-driven methods for predicting chemical properties from molecular structure. For many decades, the cheminformatics field has relied heavily on structural fingerprinting, while in recent years much focus has shifted toward leveraging highly parameterized deep neural networks which usually maximize accuracy. Beyond accuracy, to be useful and trustworthy in scientific applications, machine learning techniques often need intuitive explanations for model predictions and uncertainty quantification techniques so a practitioner might know when a model is appropriate to apply to new data. Here we revisit graphlet histogram fingerprints and introduce several new elements. We show that linear models built on graphlet fingerprints attain accuracy that is competitive with the state of the art while retaining an explainability advantage over black-box approaches. We show how to produce precise explanations of predictions by exploiting the relationships between molecular graphlets and show that these explanations are consistent with chemical intuition, experimental measurements, and theoretical calculations. Finally, we show how to use the presence of unseen fragments in new molecules to adjust predictions and quantify uncertainty.
Rare-earth and actinide complexes are critical for a wealth of clean-energy applications. Three-dimensional (3D) structural generation and prediction for these organometallic systems remains a challenge, limiting opportunities for computational chemical discovery. Here, we introduce Architector, a high-throughput in-silico synthesis code for s-, p-, d-, and f-block mononuclear organometallic complexes capable of capturing nearly the full diversity of the known experimental chemical space. Beyond known chemical space, Architector performs in-silico design of new complexes including any chemically accessible metal-ligand combinations. Architector leverages metal-center symmetry, interatomic force fields, and tight binding methods to build many possible 3D conformers from minimal 2D inputs including metal oxidation and spin state. Over a set of more than 6,000 x-ray diffraction (XRD)-determined complexes spanning the periodic table, we demonstrate quantitative agreement between Architector-predicted and experimentally observed structures. Further, we demonstrate out-of-the box conformer generation and energetic rankings of non-minimum energy conformers produced from Architector, which are critical for exploring potential energy surfaces and training force fields. Overall, Architector represents a transformative step towards cross-periodic table computational design of metal complex chemistry.
Machine learning (ML) plays a growing role in the design and discovery of chemicals, aiming to reduce the need to perform expensive experiments and simulations. ML for such applications is promising but difficult, as models must generalize to vast chemical spaces from small training sets and must have reliable uncertainty quantification metrics to identify and prioritize unexplored regions. Ab initio computational chemistry and chemical intuition alike often take advantage of differences between chemical conditions, rather than their absolute structure or state, to generate more reliable results. We have developed an analogous comparison-based approach for ML regression, called pairwise difference regression (PADRE), which is applicable to arbitrary underlying learning models and operates on pairs of input data points. During training, the model learns to predict differences between all possible pairs of input points. During prediction, the test points are paired with all training set points, giving rise to a set of predictions that can be treated as a distribution of which the mean is treated as a final prediction and the dispersion is treated as an uncertainty measure. Pairwise difference regression was shown to reliably improve the performance of the random forest algorithm across five chemical ML tasks. Additionally, the pair-derived dispersion is both well correlated with model error and performs well in active learning. We also show that this method is competitive with state-of-the-art neural network techniques. Thus, pairwise difference regression is a promising tool for candidate selection algorithms used in chemical discovery.
The Ziman formulation of electrical conductivity is tested in warm and hot dense matter using the pseudo-atom molecular dynamics method. Several implementation options that have been widely used in the literature are systematically tested through a comparison to the accurate, but expensive Kohn–Sham density functional theory molecular dynamics (KS-DFT-MD) calculations. The comparison is made for several elements and mixtures and for a wide range of temperatures and densities, and reveals a preferred method that generally gives very good agreement with the KS-DFT-MD results, but at a fraction of the computational cost.
The two primary purposes of LANL’s Computational Physics Student Summer Workshop are (1) To educate graduate and exceptional undergraduate students in the challenges and applications of computational physics of interest to LANL, and (2) Entice their interest toward those challenges. Computational physics is emerging as a discipline in its own right, combining expertise in mathematics, physics, and computer science. The mathematical aspects focus on numerical methods for solving equations on the computer as well as developing test problems with analytical solutions. The physics aspects are very broad, ranging from low-temperature material modeling to extremely high temperature plasma physics, radiation transport and neutron transport. The computer science issues are concerned with matching numerical algorithms to emerging architectures and maintaining the quality of extremely large codes built to perform multi-physics calculations. Although graduate programs associated with computational physics are emerging, it is apparent that the pool of U.S. citizens in this multi-disciplinary field is relatively small and is typically not focused on the aspects that are of primary interest to LANL. Furthermore, more structured foundations for LANL interaction with universities in computational physics is needed; historically interactions rely heavily on individuals’ personalities and personal contacts. Thus a tertiary purpose of the Summer Workshop is to build an educational network of LANL researchers, university professors, and emerging students to advance the field and LANL’s involvement in it. This report includes both the background for the program and the reports from the students.
The Ziman formulation of electrical conductivity is tested in warm and hot dense matter using the pseudo-atom molecular dynamics method. Several implementation options that have been widely used in the literature are systematically tested through a comparison to accurate but expensive Kohn-Sham density functional theory molecular dynamics (KS-DFT-MD) calculations. The comparison is made for several elements and mixtures and for a wide range of temperatures and densities, and reveals a preferred method that generally gives very good agreement with the KSDFT-MD results, but at a fraction of the computational cost.