The hydrogenation of CO2 to methanol is a key reaction for sustainable fuel synthesis. A crucial aspect of the catalytic mechanism is the role of monodentate formate (HCOOm*) in the initial steps of CO2 hydrogenation on Pd-based alloy catalysts, which we have investigated using density functional theory (DFT) together with subgroup discovery (SGD) analysis. The reactivity and stability of CO2 and formate species are examined on monometallic Pd, Cu, Zn surfaces and alloyed CuPd and PdZn surfaces. PdZn surfaces show low activation energy barriers for CO2 hydrogenation and, combined with weak CO2δ- adsorption energy, this suggests that an Eley-Rideal mechanism may dominate over Langmuir-Hinshelwood pathways. The adsorption energy of the monodentate formate intermediate is found to correlate significantly with the activation energy of CO2 hydrogenation across all investigated facets, where stronger adsorption yields lower activation energy, enabling its use as a predictive descriptor. To determine possible new catalytic materials, a dataset of 49 Pd-based single-atom alloy (SAA) surfaces is screened using SGD, identifying key electronic parameters, the dopant and site electron affinity, as drivers of strong HCOOm* adsorption. The obtained subgroup discovery rules highlight Mo, Nb, and W as promising earth-abundant dopants for Pd-based catalysts, further confirmed by DFT calculations. These findings offer a mechanistic rationale for catalyst design and demonstrate the utility of AI-guided screening in identifying efficient, sustainable alloy compositions to be used as catalysts for CO2 conversion.
Interpretable AI can reveal physical principles governing intricate materials properties by uncovering explicit relationships between physical parameters and target properties. The sure-independence screening and sparsifying operator (SISSO) symbolic-regression approach identifies analytical expressions that correlate a target property with a small set of parameters, termed materials genes, selected from a large pool of candidates. However, multiple gene combinations can yield equally accurate SISSO models, with individual genes contributing with different weights. Here, we establish a derivative-based sensitivity analysis that resolves the non-uniqueness of symbolic-regression descriptions, enhances interpretability, thereby enabling deeper physical insight. This analysis reveals how distinct gene combinations encode equivalent information and identifies valence orbital radii, nuclear charges, and their products as the key quantities governing the equilibrium lattice constant of perovskites.
Descriptors link basic physicochemical parameters that characterize the materials and the environment to the catalytic performance. Traditionally, descriptors are rooted in mechanistic understanding of elementary surface reactions gained from surface science and atomistic simulations on well-defined surfaces and under vacuum. However, real-world catalysis operates under elevated pressures and temperatures, where an intricate interplay of multiple physical processes, including significant materials' restructuring and transport phenomena, governs performance. To bridge this gap, we introduced an interpretable artificial intelligence (AI) approach that identifies key physicochemical parameters correlated with the measured catalytic performance. Analogous to genes in biology and medicine, these "materials genes" provide a statistical description of catalysis without requiring the explicit atomistic description of the underlying physical processes. Here, we combine the sure-independence-screening-and-sparsifying-operator (SISSO) symbolic-regression AI approach with a sensitivity analysis based on partial derivatives to determine the most influential genes needed to describe the selectivity of supported palladium-based metal alloy nanoparticles in the hydrogenation of concentrated acetylene streams. The identified genes include the calculated average d-band center and the measured average particle diameter, indicating the crucial role of adsorption and structure sensitivity on the formation of ethylene.
Band structure unfolding is a key technique for analyzing and simplifying the electronic band structure of large, internally distorted supercells that break the primitive cell's translational symmetry. In this work, we present an efficient band unfolding method for atomic orbital (AO) basis sets that explicitly accounts for both the nonorthogonality of atomic orbitals and their atom-centered nature. Unlike existing approaches that typically rely on a plane-wave representation of the (semi)valence states, we here derive analytical expressions that recast the primitive cell translational operator and the associated Bloch functions in the supercell AO basis. In turn, this enables the accurate and efficient unfolding of conduction, valence, and core states in all-electron codes, as demonstrated by our implementation in the all-electron ab initio simulation package FHI-AIMS, which employs numeric atom-centered orbitals. We explicitly demonstrate the capability of running large-scale unfolding calculations for systems with thousands of atoms and showcase the importance of this technique for computing temperature-dependent spectral functions in strongly anharmonic materials using CuI as example.
A series of PdZn/TiO2 catalysts prepared by chemical vapor impregnation (CVI) were tested for CO2 hydrogenation at 20 bar pressure and at temperatures of 230-270 degrees C. Changing the Pd and Zn molar ratio (Zn:Pd = 0-20) in a PdZn/TiO2 catalyst has a dramatic effect on selectivity for the CO2 hydrogenation reaction. Pd alone shows three main products: methanol, CO, and methane. Addition of small quantities of Zn results in the formation of a PdZn alloy, preventing methanation. At equimolar ratios of Pd and Zn, a 1:1 beta-PdZn alloy is formed and a reverse water gas shift catalyst is produced. Adding Zn in excess relative to the Pd loading results in the formation of ZnO on the TiO2 surface in addition to the PdZn alloy, dramatically increasing methanol selectivity from 5% at Zn:Pd = 1 to 55% for Zn:Pd = 2. Through a combination of theory and experiment, the active site for methanol synthesis is concluded to be the interface between PdZn nanoparticles and the ZnO overlayer on the TiO2, where interfacial formate can react with hydrogen dissociated by the metal nanoparticle.
Bayesian optimization (BO) efficiently explores vast design spaces using probabilistic surrogate models, enabling the guided discovery of materials with desired properties. However, most BO frameworks rely on the knowledge of a few key physical parameters (features) correlated with the materials property of interest. This is a challenge in heterogeneous catalysis, where the material properties are governed by an intricate interplay of multiple physical processes, and the mentioned parameters are typically unknown. Here, we introduce the Sparse Adaptive Representation-based Bayesian Optimization (SARBO) framework that utilizes the sure independence screening and sparsifying operator (SISSO) symbolic-regression method for on-the-fly selection of key physical parameters correlated with materials properties during BO. Crucially, SISSO takes into account nonlinear relationships and interactions between multiple parameters when selecting key features. We demonstrate that SARBO enables efficient navigation of the materials spaces and outperforms widely used feature-selection approaches for the simulated discovery of single- and dual-atom alloy surface sites capable of activating CO2, a critical step in the CO2 reduction reaction.
Predicting charge transport in strongly anharmonic materials, particularly ultralow thermal conductors, remains a major challenge for first-principles methods. In such systems, perturbative treatments of electron-phonon interactions and the harmonic phonon picture often break down, necessitating non-perturbative approaches. The ab initio Kubo-Greenwood(aiKG) formalism provides a rigorous framework for evaluating temperature-dependent carrier transport beyond the harmonic approximation. Nevertheless, its practical application is computationally demanding because it requires large supercells, extensive statistical sampling, and extrapolation to the zero-frequency limit. In this work, we introduce an artificial-intelligence(AI)-assisted aiKG framework that incorporates the deep-learning Hamiltonian model. By predicting the Kohn-Sham Hamiltonian with sub-meV accuracy for supercells of up to 250 atoms, the model bypasses the costly iterative self-consistent field calculations while retaining first-principles reliability within the scope of effects captured by the training data. Using a strongly anharmonic thermal insulator, potassium iodide(KI) as a benchmark system, we demonstrate that the proposed approach enables efficient simulations of electronic structure and transport properties from a large supercell. The framework reproduces temperature-dependent carrier mobilities, spectral functions, and effective masses in close agreement with the underlying density functional theory while reducing computational cost to 10
Artificial intelligence (AI) can accelerate materials design by identifying the key parameters correlated with the performance. However, widely used AI methods require big data, and only the smallest part of the available data in heterogeneous catalysis meets the quality requirement for data-efficient AI. Here, we use rigorous experimental procedures, designed to consistently take into account the kinetics of the catalyst active states formation, in order to measure 55 physicochemical parameters as well as the reactivity of 12 catalysts towards ethane, propane, and n-butane oxidation. These catalyst materials are based on vanadium or manganese redox-active elements (RAEs) and present diverse phase compositions, crystallinities, and catalytic behaviors. By applying the sure-independence-screening-and-sparsifying-operator (SISSO) approach to the consistent data set, we identify nonlinear property-function relationships depending on several key parameters, reflecting the intricate interplay of underlying processes governing selective oxidation. This approach indicates the most relevant characterization techniques and shows how the catalyst properties may be tuned in order to achieve the desired performance. For example, to achieve high olefin yields, the catalyst must have a high specific surface area, a low concentration of surface RAE, and the ability to change the surface RAE oxidation states under reaction conditions with respect to vacuum. These parameters are measured by N2 adsorption, x-ray photoelectron spectroscopy (XPS), and near-ambient-pressure in situ XPS. They reflect the relevance of local transport, site isolation, surface redox activity, and the materials dynamical restructuring under reaction conditions. Although the relationship describing the even more challenging oxygenate yields shares similarities with that for olefin yields, a parameter reflecting the importance of specific surface sites, derived from the analysis of the carbon 1s XPS spectra, is additionally identified as key for high selectivity to oxygenates.
Perovskites with tunable and switchable polarization hold immense promise for unlocking novel functionalities. Using density-functional theory, we reveal that intrinsic defects can induce, enhance, and control polarization in nonferroelectric perovskites, with SrTiO3 as our model system. At high defect concentrations, these systems exhibit strong spontaneous polarization-comparable to that of conventional ferroelectrics. Crucially, this polarization is switchable, enabled by the inherent symmetry-equivalence of defect sites in SrTiO3. Strikingly, polarization switching not only reverses the polarization direction and modulates its magnitude but also modifies the spatial distribution of localized defect states. This dynamic behavior points to unprecedented responses to external stimuli, opening new avenues for defect-engineered materials design.
The description of heterogeneous catalysis is challenged by the intricacy of numerous multi-scale processes that govern the performance of catalyst materials. The chemical environment of the catalytic process and the kinetics of structural changes create configurations of typically unknown local geometries and chemistry. These may result in significant changes in activity or selectivity within minutes, hours, or longer, during the so-called induction period. Here, we use experimental data together with a focused artificial-intelligence (AI) approach based on subgroup discovery and symbolic regression to model the evolution of the catalyst reactivity with time on stream. We consider palladium-based alloys synthesized mechanochemically and applied in the selective hydrogenation of concentrated acetylene streams resulting from a hypothetical electric plasma-assisted methane-to-ethylene process. Our AI approach starts with the identification of descriptions of materials and reaction conditions relevant to acetylene conversion. Then, a model for time-on-stream-dependent selectivity focused on situations associated to noticeable acetylene conversion is obtained by the sure-independence-screening-and-sparsifying-operator (SISSO) approach. Our AI approach identifies relationships between the measured catalyst reactivity and only few, key parameters, from 21 measured and calculated bulk, surface, and mesoscopic materials' properties and reaction parameters offered as candidate descriptive parameters. These identified parameters highlight the crucial influence of surface and subsurface carbon and hydrogen on the selectivity towards ethylene formation. Guided by the AI models, new, highly selective bimetallic and trimetallic systems are designed and tested experimentally.
Correction for “Atomate2: modular workflows for materials science” by Alex M. Ganose et al., Digital Discovery, 2025, 4, 1944–1973, https://doi.org/10.1039/D5DD00019J.
While the periodic equation-of-motion coupled-cluster (EOM-CC) method promises systematic improvement of electronic band gap calculations in solids, its practical application at the singles and doubles level (EOMCCSD) is hindered by severe finite-size errors in feasible simulation cells. We present a hybrid approach combining EOM-CCSD with the computationally less demanding GW approximation to estimate thermodynamic limit band gaps for several insulators and semiconductors. Our method substantially reduces required cell sizes while maintaining accuracy. Comparisons with experimental gaps and self-consistent GW calculations reveal that deviations in EOM-CCSD predictions correlate with reduced single excitation character of the excited many-electron states. Our work not only provides a computationally tractable approach to EOM-CC calculations in solids but also reveals fundamental insights into the role of single excitations in electronic-structure theory.
Sure-independence screening and sparsifying operator (SISSO) is an artificial intelligence (AI) method based on symbolic regression and compressed sensing widely used in materials science research. SISSO++ is its C++ implementation that employs MPI and OpenMP for parallelization, rendering it well-suited for high-performance computing (HPC) environments. As heterogeneous hardware becomes mainstream in the HPC and AI fields, we chose to port the SISSO++ code to GPUs using the Kokkos performance-portable library. Kokkos allows us to maintain a single codebase for both Nvidia and AMD GPUs, significantly reducing the maintenance effort. In this work, we summarize the necessary code changes we did to achieve hardware and performance portability. This is accompanied by performance benchmarks on Nvidia and AMD GPUs. We demonstrate the speedups obtained from using GPUs across the three most time-consuming parts of our code.
This paper represents one contribution to a larger Roadmap article reviewing the current status of the FHI-aims code. In this contribution, the implementation of density-functional perturbation theory in a numerical atom-centered framework is summarized. Guidelines on usage and links to tutorials are provided.
We investigate the convergence of quasiparticle energies for periodic systems to the thermodynamic limit using increasingly large simulation cells corresponding to increasingly dense integration meshes in reciprocal space. The quasiparticle energies are computed at the level of equation-of-motion coupled-cluster theory for ionization (IP-EOM-CC) and electron attachment processes (EA-EOM-CC). By introducing an electronic correlation structure factor, the expected asymptotic convergence rates for systems with different dimensionality are formally derived. We rigorously test these derivations through numerical simulations for trans-polyacetylene using IP/EA-EOM-CCSD and the G0W0@HF approximation, which confirm the predicted convergence behavior. Our findings provide a solid foundation for efficient schemes to correct finite-size errors in IP/EA-EOM-CCSD calculations.
Describing heterogeneous catalysis is complicated by the intricate interplay of processes that govern catalyst performance. The evolving chemical environment and the kinetics of catalyst's structural changes during reactions often lead to unknown local geometries and chemistry, which can shift reactivity over time. Here, we perform systematic experiments and apply a focused artificial-intelligence (AI) approach to model the measured time-on-stream-dependent reactivity of palladium-based bimetallic catalysts. These materials are synthesized via mechanochemistry and applied in the selective hydrogenation of concentrated acetylene streams(>14.0 vol %)under industrially relevant pressures (10 bar), resulting from a hypothetical electric plasma-assisted methane-to-ethylene process. Unlike the well-established hydrogenation of diluted acetylene (0.1 to 2.0 vol %) streams of naphtha steam cracking, the hydrogenation of concentrated acetylene streams remains largely underexplored due to the harsh reaction conditions and the explosive nature of acetylene. This precludes operando characterization or atomistic simulations to investigate catalyst time-on-stream behavior under realistic conditions. Our AI approach first uses subgroup discovery to identify descriptions of materials and reaction conditions resulting in noticeable acetylene conversion. Then, it models time-dependent selectivity focused on high acetylene conversion via the sure-independence-screening-and-sparsifying operator symbolic-regression approach. AI identifies key experimental and theoretical physicochemical descriptive parameters correlated with the reactivity, which highlight the critical interplay between the material structure and the chemical potential of the reaction mixture. The AI models enable the design of bimetallic and trimetallic catalysts, which are experimentally validated.
High-throughput density functional theory (DFT) calculations have become a vital element of computational materials science, enabling materials screening, property database generation, and training of "universal" machine learning models. While several software frameworks have emerged to support these computational efforts, new developments such as machine learned force fields have increased demands for more flexible and programmable workflow solutions. This manuscript introduces atomate2, a comprehensive evolution of our original atomate framework, designed to address existing limitations in computational materials research infrastructure. Key features include the support for multiple electronic structure packages and interoperability between them, along with generalizable workflows that can be written in an abstract form irrespective of the DFT package or machine learning force field used within them. Our hope is that atomate2's improved usability and extensibility can reduce technical barriers for high-throughput research workflows and facilitate the rapid adoption of emerging methods in computational material science.
Materials databases built from calculations based on density functional approximations play an important role in the discovery of materials with improved properties. Most databases thus constructed rely on the generalized gradient approximation (GGA) for electron exchange and correlation. This limits the reliability of these databases, as well as that of the artificial intelligence (AI) models trained on them, in particular for materials and properties which are not accurately described by GGA. Here, we describe a database of 7,024 inorganic materials presenting diverse structures and compositions. Crucially, the database was generated using hybrid functional calculations,efficiently implemented in the all-electron code FHI-aims. The database is used to evaluate the thermodynamic and electrochemical stability of oxides relevant to catalysis and energy related applications. We illustrate how the database can be used to train AI models for material properties using the sure-independence screening and sparsifying operator (SISSO) approach.
The International Workshop on Data-Driven Computational and Theoretical Materials Design was held between October 9-13, 2024, in Shanghai, gathering leading scientists and researchers from around the world, representing various aspects of data-driven AI methodologies and applications in materials design. The topics covered over 46 talks and 29 posters spanned a wide range of the latest advancements, including Machine Learning for Materials Design, Method Development, Machine Learning Interatomic Potentials, Advanced Computing, Infrastructure and Standards, Large Language Models, and Autonomous Labs. As part of the workshop, a panel discussion titled "Unlocking the AI Future of Materials Science" was held to disseminate the state-of-the-art of AI/ML in materials science and consider directions for the future. This report is a synthesis, for this Special Issue, of the panel discussion-drawing on insights gained from the workshop as a whole and surrounding conversations, in particular, the question of what constitutes success.
The numerical precision of density-functional-theory (DFT) calculations depends on a variety of computational parameters, one of the most critical being the basis-set size. The ultimate precision is reached in the limit of a complete basis set (CBS). Our aim in this work is to find a machine-learning model that extrapolates finite basis-size calculations to the CBS limit for periodic crystal structures. We start with a data set of 63 binary solids investigated with two all-electron DFT codes, and FHI-aims, which employ very different types of basis sets. A quantile-random-forest model and a symbolic regression approach using the SISSO model are used to estimate the total-energy correction with respect to a fully converged calculation as a function of the basis-set size. The random-forest model achieves a symmetric mean absolute percentage error of lower than 25% for both codes and outperforms previous approaches in the literature. SISSO outperforms the random forest model for the code. Our approach also provides prediction intervals, which quantify the uncertainty of the models' predictions. Published by the American Physical Society 2025