New methods for carbon dioxide removal are urgently needed to combat global climate change. Direct air capture (DAC) is an emerging technology to capture carbon dioxide directly from ambient air. Metal-organic frameworks (MOFs) have been widely studied as potentially customizable adsorbents for DAC. However, discovering promising MOF sorbents for DAC is challenging because of the vast chemical space to explore and the need to understand materials as functions of humidity and temperature. We explore a computational approach benefiting from recent innovations in machine learning (ML) and present a dataset named Open DAC 2023 (ODAC23) consisting of more than 38M density functional theory (DFT) calculations on more than 8,400 MOF materials containing adsorbed CO_2 and/or H_2O. ODAC23 is by far the largest dataset of MOF adsorption calculations at the DFT level of accuracy currently available. In addition to probing properties of adsorbed molecules, the dataset is a rich source of information on structural relaxation of MOFs, which will be useful in many contexts beyond specific applications for DAC. A large number of MOFs with promising properties for DAC are identified directly in ODAC23. We also trained state-of-the-art ML models on this dataset to approximate calculations at the DFT level. This open-source dataset and our initial ML models will provide an important baseline for future efforts to identify MOFs for a wide range of applications, including DAC.
Energy-related descriptors in machine learning are a promising strategy to predict adsorption properties of metal-organic frameworks (MOFs) in the low-pressure regime. Interactions between hosts and guests in these systems are typically expressed as a sum of dispersion and electrostatic potentials. The energy landscape of dispersion potentials plays a crucial role in defining Henry's constants for simple probe molecules in MOFs. To incorporate more information about this energy landscape, we introduce the Gaussian-approximated Lennard-Jones (GALJ) potential, which fits pairwise Lennard-Jones potentials with multiple Gaussians by varying their heights and widths. The GALJ approach is capable of replicating information that can be obtained from the original LJ potentials and enables efficient development of Gaussian integral (GI) descriptors that account for spatial correlations in the dispersion energy environment. GI descriptors would be computationally inconvenient to compute using the usual direct evaluation of the dispersion potential energy surface. We show that these new GI descriptors lead to improvement in ML predictions of Henry's constants for a diverse set of adsorbates in MOFs compared to previous approaches to this task.
Adsorption- based separations using metal- organic frameworks (MOFs) are a promising alternative to traditional energy-intensive separation process. Machine learning (ML) methods have been applied to predict large collections of adsorption isotherms in MOFs. Previous ML models, however, focus only on predicting single-component adsorption isotherms of a small number of molecules at a single temperature and lack accuracy in the dilute limit. Here we describe a useful strategy for predicting Henry's constants and heats of adsorption for a diverse set of molecules in large collections of MOFs. To achieve this, a data set containing 21,195 MOF-molecule pairs with 45 adsorbates in 471 MOFs is generated, and a set of 135 descriptors combining energy and chemical information is developed. Robust ML models are developed to predict Henry's constants and heats of adsorption after removing physically unfavorable adsorption pairs. The adsorption selectivity of near-azeotropic mixtures at two temperatures (300 and 373 K) is predicted with acceptable accuracy by using the predicted Henry's constants and heats of adsorption. The ability to make temperature-dependent predictions is important for many practical separation applications. Our work sheds light on important challenges and opportunities for developing accurate models predicting adsorption properties for diverse collection of adsorbates and adsorbents.
We have developed a simple text mining algorithm that allows us to identify surface area and pore volumes of metal-organic frameworks (MOFs) using manuscript html files as inputs. The algorithm searches for common units (e.g., m2/g, cm3/g) associated with these two quantities to facilitate the search. From the sample set data of over 200 MOFs, the algorithm managed to identify 90% and 88.8% of the correct surface area and pore volume values. Further application to a test set of randomly chosen MOF html files yielded 73.2% and 85.1% accuracies for the two respective quantities. Most of the errors stem from unorthodox sentence structures that made it difficult to identify the correct data as well as bolded notations of MOFs (e.g., 1a) that made it difficult identify its real name. These types of tools will become useful when it comes to discovering structure-property relationships among MOFs as well as collecting a large set of data for references.