Artificial Intelligence (AI) in materials science is driving significant advancements in the discovery of advanced materials for energy applications. The recent GNoME protocol identifies over 380,000 novel stable crystals. From this, we identify over 38,500 materials with potential as energy materials forming the core of the Energy-GNoME database. Our unique combination of Machine Learning (ML) and Deep Learning (DL) tools mitigates cross-domain data bias using feature spaces, thus identifying potential candidates for thermoelectric materials, novel battery cathodes, and novel perovskites. First, classifiers with both structural and compositional features detect domains of applicability, where we expect enhanced reliability of regressors. Here, regressors are trained to predict key materials properties, like thermoelectric figure of merit (zT), band gap (Eg), and cathode voltage (4V). This method significantly narrows the pool of potential candidates, serving as an efficient guide for experimental and computational chemistry investigations and accelerating the discovery of materials suited for electricity generation, energy storage and conversion.
The sunlight-driven reduction of CO2 into fuels and platform chemicals is a promising approach to enable a circular economy. However, established optimization approaches are poorly suited to multivariable multimetric photocatalytic systems because they aim to optimize one performance metric while sacrificing the others and thereby limit overall system performance. Herein, we address this multimetric challenge by defining a metric for holistic system performance that takes multiple figures of merit into account, and employ a machine learning algorithm to efficiently guide our experiments through the large parameter matrix to make holistic optimization accessible for human experimentalists. As a test platform, we employ a five-component system that self-assembles into photocatalytic micelles for CO2-to-CO reduction, which we experimentally optimized to simultaneously improve yield, quantum yield, turnover number, and frequency while maintaining high selectivity. Leveraging the data set with machine learning algorithms allows quantification of each parameter's effect on overall system performance. The buffer concentration is unexpectedly revealed as the dominating parameter for optimal photocatalytic activity, and is nearly four times more important than the catalyst concentration. The expanded use and standardization of this methodology to define and optimize holistic performance will accelerate progress in different areas of catalysis by providing unprecedented insights into performance bottlenecks, enhancing comparability, and taking results beyond comparison of subjective figures of merit.
It stands to reason that the amount and the quality of data are of key importance for setting up accurate artificial intelligence (AI)-driven models. Among others, a fundamental aspect to consider is the bias introduced during sample selection in database generation. This is particularly relevant when a model is trained on a specialized data set to predict a property of interest and then applied to forecast the same property over samples having a completely different genesis. Indeed, the resulting biased model will likely produce unreliable predictions for many of those out-of-the-box samples, i.e., samples out of the training set. Neglecting such an aspect may hinder the AI-based discovery process, even when high-quality, sufficiently large, and highly reputable data sources are available. To address this challenge, we propose a new method that detects and quantifies data bias, reducing its impact on materials discovery. Our approach, aimed at identifying and excluding those out-of-the-box materials for which the predictions of a pretrained model are likely unreliable, leverages a classification strategy and is validated by means of superconductor and thermoelectric materials as two representative case studies. This methodology, designed to be simple, flexible, and easily adaptable to any architecture, including modern graph equivariant neural networks, aims to enhance the reliability of AI models when applied to diverse and previously unseen materials, thereby contributing to more reliable AI-driven materials discovery.
We assume that a sufficiently large database is available, where a physical property of interest and a number of associated ruling primitive variables or observables are stored. We introduce and test two machine learning approaches to discover possible groups or combinations of primitive variables, regardless of data origin, being it numerical or experimental: the first approach is based on regression models, whereas the second on classification models. The variable group (here referred to as the new effective good variable) can be considered as successfully found when the physical property of interest is characterized by the following effective invariant behavior: in the first method, invariance of the group implies invariance of the property up to a given accuracy; in the other method, upon partition of the physical property values into two or more classes, invariance of the group implies invariance of the class. For the sake of illustration, the two methods are successfully applied to two popular empirical correlations describing the convective heat transfer phenomenon and to the Newton’s law of universal gravitation.
A comparison of several classifiers is presented, with a focus on the key choice and construction of a minimal set of suitable material features. To this end, an investigation is conducted over a properly selected and high quality database reporting low temperature superconductors, featurized by composition-based descriptors. Fully general strategies to reduce the number of descriptors for material classification are proposed and discussed. The first strategy aims at testing possible invariance of the target material property (here the critical temperature) with respect to (binary) groups of composition-based features in the form xiaxjb,a,b∈R. In addition, a multi-objective optimization procedure for reducing the set of composition-based material descriptors is also suggested and tested on the chosen use case. The latter procedure is then proven to be particularly convenient to be used in combination with Bayesian type classifiers. Finally, by means of the best-performing classification models, an analysis is conducted over all the ∼40,000 inorganic compounds without Ni, Fe, Cu, O in Materials Project (and not in the SuperCon database, here used for model training) and the corresponding predictions are provided. Among those, 41 materials are classified to show Tc≥15K with a probability higher than or equal to 0.6.
It stands to reason that the amount and the quality of big data is of key importance for setting up accurate AI-driven models. Nonetheless, we believe there are still critical roadblocks in the inherent generation of databases, that are often underestimated and poorly discussed in the literature. In our view, such issues can seriously hinder the AI-based discovery process, even when high quality, sufficiently large and highly reputable data sources are available. Here, considering superconducting and thermoelectric materials as two representative case studies, we specifically discuss three aspects, namely intrinsically biased sample selection, possible hidden variables, disparate data age. Importantly, to our knowledge, we suggest and test a first strategy capable of detecting and quantifying the presence of the intrinsic data bias.
In this study, we evaluate several classifiers and focus on selecting a minimal set of appropriate material features. Our objective is to propose and discuss general strategies for reducing the number of descriptors required for material classification. The first strategy involves testing whether the critical temperature of the target material property is invariant with respect to binary groups of composition-based features. We also propose a multi-objective optimization procedure to reduce the set of composition-based material descriptors. The latter procedure is found to be particularly useful when applied to Bayesian classifiers. We test the proposed strategies focusing on low-temperature superconductors material data extracted from a public database.
We focus on gas sorption within metal-organic frameworks (MOFs) for energy applications and identify the minimal set of crystallographic descriptors underpinning the most important properties of MOFs for CO2 and H2O. A comprehensive comparison of several sequential learning algorithms for MOFs properties optimization is performed and the role played by those descriptors is clarified. In energy transformations, thermodynamic limits of important figures of merit crucially depend on equilibrium properties in a wide range of sorbate coverage values, which is often only partially accessible, hence possibly preventing the computation of desired objective functions. We propose a fast procedure for optimizing specific energy in a closed sorption energy storage system with only access to a single water Henry coefficient value and to the specific surface area. We are thus able to identify hypothetical candidate MOFs that are predicted to outperform state-of-the-art water-sorbent pairs for thermal energy storage applications.
Several studies have been recently reported in the literature on sorption properties of MOFs with a number of organic sorbates, such as ethanol and methanol. Surprisingly, still few studies have been reported on water sorbate despite its large availability, low cost and environmental sustainability, and the screening of a large number of hypothetical MOFs-water working pairs for engineering applications is still challenging. Based on a recently reported database of over 5000 hypothetical MOFs, a first contribution of this study is the identification of the minimal set of crystallographic descriptors underpinning the most important sorption properties of MOFs for \ch{CO2} and, importantly, for \ch{H2O}. Furthermore, a comprehensive comparison of several Sequential Learning (SL) algorithms for MOFs properties optimization is carried out and the role played by the above minimal set of crystallographic descriptors clarified. In sorption-based energy transformations, thermodynamic limits of important figures of merit (e.g. maximum specific energy) depend both on operating conditions and equilibrium sorption properties in a wide range of sorbate coverage values. The access to the latter properties is often incomplete, with essential quantities such as equilibrium adsorption isotherms spanning over the full sorbate coverage range and values of the isosteric heat being only partially available. As a result, this may prevent the computation of objective functions during the optimization procedure. We propose a fast procedure for optimizing specific energy in a closed sorption energy storage system with the only access to the water Henry coefficient at a fixed temperature value and to the specific surface area.