The Active Thermochemical Tables (ATcT) methodology represents a paradigm shift from traditional sequential thermochemistry to a network-based framework in which all available experimental and theoretical determinations are simultaneously incorporated and statistically reconciled within an overdetermined thermochemical network (TN). This perspective examines the role of ATcT as critical data infrastructure for theory development, method validation, and the emerging integration of artificial intelligence in the chemical sciences. We describe the structural limitations of legacy sequential thermochemical compilations and contrast them with the ATcT approach, which yields internally consistent thermochemical values with rigorously quantified uncertainties, full covariance structure, and quantitative provenance through variance decomposition. We present the first explicit formulation of edgewise uncertainty decomposition and leverage diagnostics within the context of thermochemical networks. The development of a RESTful API and an open-source Python client (atct) is described, providing machine-actionable, FAIR-compliant, and versioned programmatic access to ATcT data─including species correlations and covariance-aware reaction uncertainty propagation─capabilities essential for high-throughput benchmarking and automated computational workflows. We discuss the critical distinction between mean absolute deviation and 95% confidence intervals in method assessment, and the implications of using curated versus aggregated data for training machine learning models. The susceptibility of large language models to thermochemical hallucination is illustrated through the instructive case of the electron affinity of BH3, underscoring the necessity of coupling AI systems to authoritative, uncertainty-quantified data sources. ATcT's designation as a U.S. Department of Energy Office of Science Public Reusable Research (DOE SC PuRe) Data Resource ensures its long-term stewardship as foundational infrastructure for both traditional computational thermochemistry and next-generation AI-driven chemical research.
The reaction of CH + N2 forming H + NCN is a remarkable example of activation of the nitrogen triple bond and is an important source of prompt NO in combustion. The reaction pathway is complex and proceeds through two competing mechanisms: a cyclic addition channel initiated by c-HC(NN) and a chain-addition channel initiated by HCNN, both of which eventually form HNCN prior to dissociation to H + NCN. This work reinvestigates this reaction with composite coupled cluster protocols, including a novel spin-flip equation of motion coupled cluster scheme, combined with pragmatic two-dimensional master equation simulations of the resulting rate coefficients. These improved calculations predict the CH + N2 rate coefficient between the two more recent previous theoretical results and reduce the uncertainties of the best theoretical models of this reaction to less than a factor of 1.3. Additionally, we provide a closer theoretical investigation of the simultaneous dependence of the CH + N2 rate coefficient on pressure and temperature, and affirm that collisionally stabilized HNCN, another potential source of prompt NO, emerges as an appreciable product of this reaction under conditions relevant to automotive internal combustion engines and aircraft gas turbine engines.
Heterogeneous catalysis is critical in most industrial chemical processes. Microkinetic models can be used to greatly facilitate optimization of catalyst design and process conditions, but require thermochemical and kinetic parameters for all relevant species 1 and reactions. Our software Pynta enables fully automated calculation of thermochemical and kinetic parameters, however, the computational cost of density functional theory (DFT) calculations makes it difficult to calculate kinetic parameters at scale. In this work, we combine finetuning of the MACE multi-head v0 model with graph-based delta learning of stationary points using subgraph isomorphic decision trees (SIDT). We first generate a target chemical space on Pt111 by using the Reaction Mechanism Generator (RMG) software to generate all possible surface reactions between a set of 153 chemically adsorbed species containing C,H,N, and O. Applying Pynta to this reaction set we are able to generate geometries for many gas phase species, adsorbates and transition states in this chemical space allowing us to efficiently sample near stationary points for finetuning and stationary points for delta learning. In particular, we adapt Pynta's harmonically force saddle point search (HFSP) algorithm to enable efficient and reliable sampling of near transition state points. With this training data we show that SIDT-driven delta learning of the foundation MLIP alone can achieve similar accuracies on enthalpies to foundation model finetuning approaches. However, we show that by combining the two approaches into one framework we are able to significantly improve our accuracies over either approach allowing us to achieve 0.07 eV MAE relative to DFT on test barrier heights. Applying this framework across all reactions in our chemical space, we are able to obtain near DFT accuracy rate coefficients for 3111 reactions on Pt(111).
A series of approximations to CCSD contributions in computational model chemistries is presented in the context of kcal mol-1, kJ mol-1, and 20 cm-1 theoretical predictions of total atomization energies, benchmarked within the HEAT+CH4 test suite. A specific set of circumstances where MP2, without empirical scaling, may be used as an effective intermediate in the first two of these accuracy ranges was determined. However, SDQ-MP4, a method long used in pursuit of kcal mol-1 accuracy but relatively unstudied in the subchemical accuracy community, offers significant improvement over the quality of MP2 as a basis-set intermediate at significantly reduced cost compared to CCSD. Given this, we argue for SDQ-MP4 as the de facto CCSD basis-set intermediate in sub-chemical accuracy calculations when CCSD in a desired basis set becomes unaffordable. We additionally report on a "CBS-like" scheme, where MP2 and SDQ-MP4 are used in conjunction to create a "cheap" three-part approximation of large CCSD basis set limits. The data for the CCSD approximation schemes are organized in such a way that model chemistry developers can locate an analog of their current approach for the CCSD basis set limit and explore alternative intermediates that either decrease computational cost or increase computational accuracy. We also show, for a handful of molecules, that SDQ-MP4 shows promise as an effective basis-set intermediate for harmonic and fundamental frequency computations, allowing for zero-point corrections of nearly CCSD(T)/ANO1 quality using simple composite methods that only require CCSD(T)/ANO0.
High-accuracy ab initio thermochemical predictions for the ionization energy of NF3, the barrier height (to inversion) of NF3+, and the dissociative ionization threshold of NF3 to NF2+ + F are presented and incorporated into Active Thermochemical Tables. The adiabatic ionization energy of the first ionization band of NF3, calculated at 12.647 ± 0.010 eV, is at odds with previous experimental interpretations by nearly 0.36 eV due to unfavorable Franck-Condon factors associated with this transition. The barrier (to inversion) height is calculated to be about 0.6 eV lower in energy than the prior interpretation, which instigates a discussion of the supposed vibrational structure of the first ionization band of NF3. Updated assignments of the photoelectron spectrum are proposed, and the loss in vibrational spacing on the high-energy side of the experimental ionization band is discussed. Rudimentary anharmonic Franck-Condon simulations qualitatively reproduce the broad spectral features observed in experiment.
Thermophysical properties of adsorbates and gas-phase species define the free energy landscape of heterogeneously catalyzed processes and are pivotal for an atomistic understanding of the catalyst performance. These thermophysical properties, such as the free energy or the enthalpy, are typically derived from density functional theory (DFT) calculations. Enthalpies are species-interdependent properties that are only meaningful when referenced to other species. The widespread use of DFT has led to a proliferation of new energetic data in the literature and databases. However, there is a lack of consistency in how DFT data is referenced and how the associated enthalpies or free energies are stored and reported, leading to challenges in reproducing or utilizing the results of prior work. Additionally, DFT suffers from exchange-correlation errors that often require corrections to align the data with other global thermochemical networks, which are not always clearly documented or explained. In this review, we introduce a set of consistent terminology and definitions, review existing approaches, and unify the techniques using the framework of linear algebra. This set of terminology and tools facilitates the correction and alignment of energies between different data formats and sources, promoting the sharing and reuse of ab initio data. Standardization of thermochemistry concepts in computational heterogeneous catalysis reduces computational cost and enhances fundamental understanding of catalytic processes, which will accelerate the computational design of optimally performing catalysts.
In contrast to the adage "Models are to be used, not believed", combustion kinetics models have been intended to be predictive in nature. Theoretical chemical kinetics is now understood to provide a firm foundation for the reaction parameters, thereby facilitating predictive simulations of chemical reactivity, even in regimes that are poorly characterized by chemical kinetic and/or combustion experiments. We describe here a theory-informed kinetics model (ThInK) for small molecule combustion chemistry (H2 and C1 - C3 species) that is based on the prodigious use of theoretical predictions for reaction rate coefficients, thermochemistry, and transport parameters. The distinct features of this kinetics model, which was developed over the course of several decades, are illustrated through simulations of flame propagation and auto-ignition. Novelty and significance statement: The novelty of the "ThInK" C0-C3 mechanism is that an overwhelming number of its parameters are derived from a priori theoretical predictions. It marks a significant departure from traditional models that rely heavily on experimental data and adjusted or empirical parameters. This advancement represents a transformative step in the modeling of combustion kinetics, providing an exceptionally robust smallmolecule core mechanism upon which larger models can be based.
The Active Thermochemical Tables approach produces the enthalpy of formation of gas phase boron atom: O f H degrees 298 (B (g) ) = 570.48 f 0.61 kJ/mol and O f H degrees 0 (B (g) ) = 565.38 f 0.61 kJ/mol. This is about 5 kJ/mol higher and nearly an order of magnitude more accurate than the CODATA value. While the ATcT value is in excellent agreement with the revisions proposed by Bauschlicher, Martin, and Taylor [J. Phys. Chem. A 103 (1999) 7715] and by Karton and Martin [J. Phys. Chem. A 111 (2007) 5936], it invalidates several earlier theoretical revisions that are too high by up to 5 kJ/mol.
Microkinetic models for catalytic systems require estimation of many thermodynamic and kinetic parameters that can be calculated for isolated species and transition states using ab initio methods. However, the presence of nearby co-adsorbates on the surface can dramatically alter these thermodynamic and kinetic parameters causing them to be dependent on species coverage fractions. As there are combinatorially many co-adsorbed configurations on the surface, computing the coverage dependence of these parameters is far less straightforward. We present a framework for generating and applying machine learning models to predict coverage-dependent parameters for microkinetic models. Our toolkit enables automatic calculation and evaluation of co-adsorbed configurations allowing us to sample 2,000 co-adsorbed adsorbates and transition states (TSs) for a diverse set of 9 reactions on Cu(111), a challenging surface, with four possible co-adsorbates. This dataset was then used to train subgraph isomorphic decision trees (SIDTs) to predict the stability and association energy of configurations. We were able to achieve mean absolute errors (MAEs) of 0.106 eV on adsorbates, 0.172 eV on TSs, and due to natural error cancellation in SIDTs for relative properties, 0.130 eV on reaction energies and 0.180 eV on activation barriers. We describe how to use these models to predict coverage- dependent corrections for adsorbates and TSs, and demonstrate on H∗, HO∗ and O∗ comparing the generated SIDT model with an iteratively refined version.
Active Thermochemical Tables (ATcT) were successfully used to resolve the existing inconsistencies related to the thermochemistry of glycine, based on statistically analyzing and solving a thermochemical network that includes >3350 chemical species interconnected by nearly 35 000 thermochemically-relevant determinations from experiment and high-level theory. The current ATcT results for the 298.15 K enthalpies of formation are -394.70 ± 0.55 kJ mol-1 for gas phase glycine, -528.37 ± 0.20 kJ mol-1 for solid α-glycine, -528.05 ± 0.22 kJ mol-1 for β-glycine, -528.64 ± 0.23 kJ mol-1 for γ-glycine, -514.22 ± 0.20 kJ mol-1 for aqueous undissociated glycine, and -470.09 ± 0.20 kJ mol-1 for fully dissociated aqueous glycine at infinite dilution. In addition, a new set of thermophysical properties of gas phase glycine was obtained from a fully corrected nonrigid rotor anharmonic oscillator (NRRAO) partition function, which includes all conformers. Corresponding sets of thermophysical properties of α-, β-, and γ-glycine are also presented.
A new strategy is presented for computing anharmonic partition functions for the motion of adsorbates relative to a catalytic surface. Importance sampling is compared with conventional Monte Carlo. The importance sampling is significantly more efficient. This new approach is applied to CH3* on Ni(111) as a test case. The motion of methyl relative to the nickel surface is found to be anharmonic, with significantly higher entropy compared to the standard harmonic oscillator model. The new method is freely available as part of the Minima-Preserving Neural Network within the AdTherm package.
Modern plane-wave DFT methods and software (contained in the NWChem and NWChemEx packages) that allow for both geometry optimization and ab initio molecular dynamics simulations (AIMD) are described. Significant emphasis is placed on aspects of these methods that are of interest to computational chemists and useful for simulating chemistry, including techniques for calculating charged systems, exact exchange (i.e., hybrid DFT methods), and highly efficient AIMD/MM methods. Sample applications for the hydrolysis of nitroaromatic molecules, the structure of the goethite+water interface, and AIMD-EXAFS calculations for uranium metal impurities in iron-(oxyhydr)oxides are described.
Atmospheric formic acid is severely underpredicted by models. A recent study proposed that this discrepancy can be resolved by abundant formic acid production from the reaction ( 1 ) between hydroxyl radical and methanediol derived from in-cloud formaldehyde processing and provided a chamber-experiment-derived rate constant, k 1 = 7.5 × 10 −12 cm 3 s −1 . High-level accuracy coupled cluster calculations in combination with E,J -resolved two-dimensional master equation analyses yield k 1 = (2.4 ± 0.5) × 10 −12 cm 3 s −1 for relevant atmospheric conditions ( T = 260–310 K and P = 0–1 atm). We attribute this significant discrepancy to HCOOH formation from other molecules in the chamber experiments. More importantly, we show that reversible aqueous processes result indirectly in the equilibration on a 10 min. time scale of the gas-phase reaction HCHO + H 2 O ⇌ HOCH 2 OH (2) with a HOCH 2 OH to HCHO ratio of only ca . 2%. Although HOCH 2 OH outgassing upon cloud evaporation typically increases this ratio by a factor of 1.5–5, as determined by numerical simulations, its in-cloud reprocessing is shown using a global model to strongly limit the gas-phase sink and the resulting production of formic acid. Based on the combined findings in this work, we derive a range of 1.2–8.5 Tg/y for the global HCOOH production from cloud-derived HOCH 2 OH reacting with OH. The best estimate, 3.3 Tg/y, is about 30 times less than recently reported. The theoretical equilibrium constant K eq (2) determined in this work also allows us to estimate the Henry’s law constant of methanediol (8.1 × 10 5 M atm −1 at 280 K).
Many important industrial processes rely on heterogeneous catalytic systems. However, given all possible catalysts and conditions of interest, it is impractical to optimize most systems experimentally. Automatically generated microkinetic models can be used to efficiently consider many catalysts and conditions. However, these microkinetic models require accurate estimation of many thermochemical and kinetic parameters. Manually calculating these parameters is tedious and error prone, involving many interconnected computations. We present Pynta, a workflow software for automating the calculation of surface and gas-surface reactions. Pynta takes the reactants, products, and atom maps for the reactions of interest, generates sets of initial guesses for all species and saddle points, runs all optimizations, frequency, and IRC calculations, and computes the associated thermochemistry and rate coefficients. It is able to consider all unique adsorption configurations for both adsorbates and saddle points, allowing it to handle high index surfaces and bidentate species. Pynta implements a new saddle point guess generation method called harmonically forced saddle point searching (HFSP). HFSP defines harmonic potentials based on the optimized adsorbate geometries and which bonds are breaking and forming that allow initial placements to be optimized using the GFN1-xTB semiempirical method to create reliable saddle point guesses. This method is reaction class agnostic and fast, allowing Pynta to consider all possible adsorbate site placements efficiently. We demonstrate Pynta on 11 diverse reactions involving monodenate, bidentate, and gas-phase species, many distinct reaction classes, and both a low and a high index facet of Cu. Our results suggest that it is very important to consider reactions between adsorbates adsorbed in all unique configurations for interadsorbate group transfers and reactions on high index surfaces.
The bond dissociation energy of methylidyne, D0(CH), is studied using an improved version of the High-Accuracy Extrapolated ab initio Thermochemistry (HEAT) approach as well as the Feller-Peterson-Dixon (FPD) model chemistry. These calculations, which include basis sets up to nonuple (aug-cc-pCV9Z) quality, are expected to be capable of providing results substantially more accurate than the ca. 1 kJ mol-1 level that is characteristic of standard high-accuracy protocols for computational thermochemistry. The calculated 0 K CH bond energy (27 954 ± 15 cm-1 for HEAT and 27 956 ± 15 cm-1 for FPD), along with equivalent treatments of the CH ionization energy and the CH+ dissociation energy (85 829 ± 15 cm-1 and 32 946 ± 15 cm-1, respectively), were compared to the existing benchmarks from Active Thermochemical Tables (ATcT), uncovering an unexpected difference for D0(CH). This has prompted a detailed reexamination of the provenance of the corresponding ATcT benchmark, allowing the discovery and subsequent correction of a systematic error present in several published high-level calculations, ultimately yielding an amended ATcT benchmark for D0(CH). Finally, the current theoretical results were added to the ATcT Thermochemical Network, producing refined ATcT estimates of 27 957.3 ± 6.0 cm-1 for D0(CH), 32 946.7 ± 0.6 cm-1 for D0(CH+), and 85 831.0 ± 6.0 cm-1 for IE(CH).
A method for computing anharmonic thermophysical properties for adsorbates on metal surfaces has been extended to include libration, or frustrated rotation. Classical phase space integration is used with Monte Carlo sampling of the configuration space to obtain the partition function of CO on Pt(111) and CH3OH on Cu(111). A minima-preserving neural network potential energy surrogate is used within the integration routines. Direct state counting using discrete variable representation is used to benchmark the results. We find that the phase space integration approach is in excellent agreement with the direct state counting results. Comparison with standard models such as the harmonic oscillator indicates that anharmonicity contributes significantly to the thermodynamic properties of CH3OH on Cu(111). We find that there is also a considerable difference between the harmonic oscillator and phase space integration for CO on Pt(111), although the discrepancy can largely be attributed to the presence of multiple binding sites within the unit cell. We demonstrate that a multisite harmonic oscillator model might be sufficient for CO-Pt(111). A more thorough description of the potential energy surface, which can be achieved with phase space integration, is necessary for weakly bound adsorbates such as CH3OH. The thermophysical properties were used to calculate free energies of adsorption on the respective metals, and subsequently the equilibrium constants and Langmuir isotherms in relevant temperature ranges. The results show that the choice of model to obtain partition functions greatly affects the resulting surface coverages in kinetic models.
The thermochemistry of halocarbon species containing iodine and bromine is examined through an extensive interplay between new Feller-Peterson-Dixon (FPD) style composite methods and a detailed analysis of all available experimental and theoretical determinations using the thermochemical network that underlies the Active Thermochemical Tables (ATcT). From the computational viewpoint, a slower convergence of the components of composite thermochemistry methods is observed relative to species that solely contain first row elements, leading to a higher computational expense for achieving comparable levels of accuracy. Potential systematic sources of computational uncertainty are investigated, and, not surprisingly, spin-orbit coupling is found to be a critical component, particularly for iodine containing molecular species. The ATcT analysis of available experimental and theoretical determinations indicates that prior theoretical determinations have significantly larger uncertainties than originally reported, particularly in cases where molecular spin-orbit effects were ignored. Accurate and reliable heats of formation are reported for 38 halogen containing systems, based on combining the current computations with previous experimental and theoretical work via the ATcT approach.
High-level coupled cluster theory, in conjunction with Active Thermochemical Tables (ATcT) and E,J-resolved master equation calculations, was used in a study of the title reactions, which play an important role in the combustion of hydrocarbons. In the set of radical/radical reactions leading to soot formation in flames, the addition of H-atoms to alkenes is likely a common reaction, triggering the isomerization of complex hydrocarbons to aromatics. The heats of formation of C2H3, C2H4, and C2H5 are established to be 301.26 ± 0.30 at 0 K (297.22 ± 0.30 at 298 K), 60.89 ± 0.11 (52.38 ± 0.11), and 131.38 ± 0.22 (120.63 ± 0.22) kJ mol-1, respectively. The calculated rate constants from first principles agree well with experiments where they are available. Under conditions typical of high temperature combustion - where experimental work is very challenging with a consequent dearth of accurate data - we provide high-level theoretical results for kinetic modeling.