Developing new separation technologies for rare-earth elements is essential for sustaining the critical materials supply chain. Toward this end, we developed a pH-controlled solvent extraction strategy employing an aqueous-phase holdback agent to enable selective lanthanide separations, using an integrated computational, machine learning, and automated experimental high-throughput workflow. Database screening, density functional theory (DFT) calculations, and initial experimental evaluation identified oxaloacetic acid as a promising holdback agent that enhanced selective extraction of four lanthanides (Nd, Eu, Dy, Ho) when paired with the di(2-ethylhexyl)phosphoric acid (HDEHP or D2EHPA) extractant. Within this framework, the dependence of lanthanide solvent extraction was determined across a multidimensional chemical matrix (pH, extractant concentration, holdback agent concentration, and salt concentration) using automated, high-throughput experiments coupled with multi-objective Bayesian Optimization. Through efficient exploration of a large experimental space, we discovered Eu, Dy, and Ho could be selectively extracted over Nd in acidic media at pH ∼2.0, while a modest decrease in pH to ∼0.5 shifted the selectivity to enable Eu separation from Dy and Ho. The use of multi-objective Bayesian Optimization quickly yielded a 4-fold increase in separation factors compared to our HDEHP-only system, eliminating the need for a complete grid-based, exhaustive sampling approach. Overall, this work establishes a hierarchical and data-driven framework to identify selective separation conditions and provides a foundation for accelerating separations discovery and design.
Advances in machine learning have given rise to a plurality of data-driven methods for predicting chemical properties from molecular structure. For many decades, the cheminformatics field has relied heavily on structural fingerprinting, while in recent years much focus has shifted toward leveraging highly parameterized deep neural networks which usually maximize accuracy. Beyond accuracy, to be useful and trustworthy in scientific applications, machine learning techniques often need intuitive explanations for model predictions and uncertainty quantification techniques so a practitioner might know when a model is appropriate to apply to new data. Here we revisit graphlet histogram fingerprints and introduce several new elements. We show that linear models built on graphlet fingerprints attain accuracy that is competitive with the state of the art while retaining an explainability advantage over black-box approaches. We show how to produce precise explanations of predictions by exploiting the relationships between molecular graphlets and show that these explanations are consistent with chemical intuition, experimental measurements, and theoretical calculations. Finally, we show how to use the presence of unseen fragments in new molecules to adjust predictions and quantify uncertainty.
Determining the activity series of a collection of elements is a classic pedagogical experiment, in which pairs of elements are reacted to determine the relative rank ordering of their reactivity. Determining the optimal sequence of pairwise experiments that minimizes the total number of experiments corresponds to well-known comparison sorting algorithms in computer science. We describe relevant algorithms (insertion sort, binary insertion sort, merge sort, and merge insertion sort) and their application to the activity series problem and discuss ways that this connection can contribute to the introductory chemistry and computer science curricula. In addition to pedagogical interest, this illustrates a simple form of artificial intelligence for chemical experiment planning.
Discrimination is a form of chronic stress and hair cortisol concentration is an emerging biomarker of chronic stress. In a sample of 83 first-year college students (age x⋅⋅−=17.65, SD=48, 69% female, 84% United States-born, 24% Asian, 21% Latinx, and 55% White), the current study investigates associations between hair cortisol concentration with discrimination stress assessed across two timeframes: past year and past two weeks. Significant associations were observed for past year discrimination and hair cortisol concentration levels, but not for discrimination over the past two weeks. The current study contributes to a growing body of evidence linking discrimination stress exposure to neuroendocrine functioning.
Machine learning (ML) plays a growing role in the design and discovery of chemicals, aiming to reduce the need to perform expensive experiments and simulations. ML for such applications is promising but difficult, as models must generalize to vast chemical spaces from small training sets and must have reliable uncertainty quantification metrics to identify and prioritize unexplored regions. Ab initio computational chemistry and chemical intuition alike often take advantage of differences between chemical conditions, rather than their absolute structure or state, to generate more reliable results. We have developed an analogous comparison-based approach for ML regression, called pairwise difference regression (PADRE), which is applicable to arbitrary underlying learning models and operates on pairs of input data points. During training, the model learns to predict differences between all possible pairs of input points. During prediction, the test points are paired with all training set points, giving rise to a set of predictions that can be treated as a distribution of which the mean is treated as a final prediction and the dispersion is treated as an uncertainty measure. Pairwise difference regression was shown to reliably improve the performance of the random forest algorithm across five chemical ML tasks. Additionally, the pair-derived dispersion is both well correlated with model error and performs well in active learning. We also show that this method is competitive with state-of-the-art neural network techniques. Thus, pairwise difference regression is a promising tool for candidate selection algorithms used in chemical discovery.
Discovery of new perovskite materials is motivated by a broad range of materials applications and accelerated by recent advances in machine learning (ML). We herein report dataset augmentation, benchmarking, and interrogation for an ongoing experimental campaign consisting of 9483 halide perovskite synthesis experiments. To address limitations in previous work, we developed an improved description of the reactant concentrations in the experiments (validated against experimental observations) and performed experiments quantifying the excess volume of mixing of gamma-butyrolactone/formic acid mixtures used in the perovskite syntheses. Combining this improved description of reactant concentration with other physicochemical features of the reactants, we constructed 1108 ML models to elucidate the roles of the algorithm (k-nearest neighbors, linear support-vector machine, and gradient boosted tree), feature set (12 in total), preprocessing regime (e.g., standardization), and training data holdout scheme on ML predictive ability. ML comparisons illustrated that the chemical accuracy of less sophisticated physical models in a dataset do not hinder interpolative model performance. Analysis of feature contributions showed how ML models "learn" competitive representations for concentration using raw experimental descriptions. Interrogation of the most performant models indicated that the numerical values of physicochemical features were not important, rather these features were being used to identify and interpolate within a particular reactant set. ML models were shown to be capable of making rudimentary extrapolations to untrained chemical systems when compared against basic benchmarks, and models which included the newly developed chemical features were shown to be more reliable than models trained without. These results illustrate how a stepwise comparative approach to machine learning can provide insight into what and how much models are "learning" for a given prediction task.
An increasing number of companies are using data analytics to improve their products, services, and business processes. However, learning knowledge effectively from massive data sets always involves nontrivial computational resources. Most businesses thus choose to migrate their hardware needs to a remote cluster computing service (e.g., AWS) or to an in-house cluster facility which is often run at its resource capacity. In such scenarios, where jobs compete for available resources utilizing resources effectively to achieve high-performance data analytics becomes desirable. Although cluster resource management is a fruitful research area having made many advances (e.g., YARN, Kubernetes), few projects have investigated how further optimizations can be made specifically for training multiple machine learning (ML) / deep learning (DL) models. In this work, we introduce FlowCon, a system which is able to monitor loss functions of ML/DL jobs at runtime, and thus to make decisions on resource configuration elastically. We present a detailed design and implementation of FlowCon, and conduct intensive experiments over various DL models. Our experimental results show that FlowCon can strongly improve DL job completion time and resource utilization efficiency, compared to existing approaches. Specifically, FlowCon can reduce the completion time by up to 42.06% for a specific job without sacrificing the overall makespan, in the presence of various DL job workloads.
Current software tools for the automated building of models for macromolecular X-ray crystal structures are capable of assembling high-quality models for ordered macromolecule and small-molecule scattering components with minimal or no user supervision. Many of these tools also incorporate robust functionality for modelling the ordered water molecules that are found in nearly all macromolecular crystal structures. However, no current tools focus on differentiating these ubiquitous water molecules from other frequently occurring multi-atom solvent species, such as sulfate, or the automated building of models for such species. PeakProbe has been developed specifically to address the need for such a tool. PeakProbe predicts likely solvent models for a given point (termed a `peak') in a structure based on analysis (`probing') of its local electron density and chemical environment. PeakProbe maps a total of 19 resolution-dependent features associated with electron density and two associated with the local chemical environment to a two-dimensional score space that is independent of resolution. Peaks are classified based on the relative frequencies with which four different classes of solvent (including water) are observed within a given region of this score space as determined by large-scale sampling of solvent models in the Protein Data Bank. Designed to classify peaks generated from difference density maxima, PeakProbe also incorporates functionality for identifying peaks associated with model errors or clusters of peaks likely to correspond to multi-atom solvent, and for the validation of existing solvent models using solvent-omit electron-density maps. When tasked with classifying peaks into one of four distinct solvent classes, PeakProbe achieves greater than 99% accuracy for both peaks derived directly from the atomic coordinates of existing solvent models and those based on difference density maxima. While the program is still under development, a fully functional version is publicly available. PeakProbe makes extensive use of cctbx libraries, and requires a PHENIX licence and an up-to-date phenix.python environment for execution.