Conventional molecular graphs are often unable to reliably encode stereochemistry, especially for symmetric molecules, nontetrahedral centers, and transition states. To overcome this, we present StereoMolGraph, an open-source Python library implementing a stereochemistry-aware graph representation for molecules and condensed graphs of reactions. Our method uses permutation-invariant local stereodescriptors, grounded in group theory, to provide an extensible representation of chirality. Based on this we introduce methods allowing for robust comparison of stereoisomers, including the identification of enantiomerism and diastereomerism, and support fleeting stereochemistry in transition states. We demonstrate the library's utility for complex organic molecules and metal complexes and analysis of distinct chiral reaction pathways. With RDKit interoperability and visualization features, StereoMolGraph offers a practical and transparent tool for advanced stereochemically aware chemoinformatics workflows.
The kinetics of the C 3 H 3 + HO 2 reaction system were investigated theoretically using high-level quantum chemical methods combined with statistical rate theory. The potential energy surface (PES) was explored to identify intermediates and transition states using reactive molecular dynamics with ChemTraYzer-TAD. The resulting thermochemical data and energy barriers for products P0 to P13 were used in RRKM/master equation calculations to determine the temperature-and pressure-dependent rate constants suitable for incorporation into chemical kinetic models. Implications for chemical kinetic model predictions and future model development were assessed by representative jet-stirred reactor simulations of stoichiometric propyne/air mixtures at 10 atm and 800-1200 K. These simulations show that the widely employed allyl + HO 2 analogy overestimates propargyl consumption. In the jet-stirred reactor case considered here, replacing the analogy-based treatment with the theory-based update reduces the contribution of C 3 H 3 + HO 2 reactions to the overall consumption of C 3 H 3 from up to 26 % to no more than 15 %. The corresponding changes in benzene predictions are visible, but remain within the experimental uncertainty. Based on these findings, a compact reduced three-reaction submodel is proposed for future chemical kinetic model development. Novelty and significance statement We present the first comprehensive ab initio and master equation-based kinetic investigation of the C 3 H 3 + HO 2 reaction system. Barrierless and new low-barrier channels were identified using potential energy surface exploration with ChemTraYzer-TAD, and the resulting temperature-and pressuredependent rate coefficients provide a substantially more reliable description of this chemistry than the allyl + HO 2 analogy commonly used in chemical kinetic models. Jet-stirred reactor simulations show that replacing the analogy-based treatment reduces predicted C 3 H 3 consumption via reactions with HO 2 and leads to visible changes in benzene predictions that remain within the experimental uncertainty. Brute force sensitivity analysis further shows that only the OH-forming channel, C 3 H 3 + HO 2 = P-C 3 H 3 O + OH, noticeably affects the predicted C 3 H 3 and benzene mole fractions. Based on these findings, a compact reduced three-reaction submodel is proposed for future chemical kinetic model development. Because C 3 H 3 is a key C0-C4 core species, the refined kinetics are relevant beyond aromatic chemistry.
Methyl ketones have recently been identified as a promising candidate for biofuels among other uses. They can be produced via biomass fermentation using Pseudomonas taiwanensis VLB120 ∆6 pProd. However, the literature does not contain a complete early stage design proposal including upstream bioreactor design and downstream product purification at industrial scale. We present a model-based design approach for methyl ketone production based on renewable resources, including fermentation and product purification. We apply OptKnock [Burgard, Pharkya, Maranas: Biotechnology and bioengineering 2003] to propose beneficial gene knockouts in the strain and adapt our SimulKnockReactor methodology [Ziegler et int. Mitsos: Comp&ChemEng 2026] to optimize the bioreactor for the production of methyl ketones. We then apply an enhanced version of our COSMO-CAMPED framework [Polte et int. Leonhard: CIT 2023] for the downstream purification of the fermentation broth. Extraction-based separations offer substantial potential for efficient purification of chemical products, particularly from the dilute aqueous streams of biotechnological fermentations. Finding a solvent with the required thermodynamic properties, bio-compatibility, and benign environmental properties is vital for the overall process performance. We propose three gene deletions to P. taiwanensis VLB120 ∆6 pProd, namely, inorganic diphosphatase, polyphosphate kinase, and inorganic diphosphatase. We identify isopentyl-isopropylether as most suitable extraction solvent among several others that outperform our literature benchmark of Di-n-butylether. We present an optimized layout for the combined upstream and downstream process that produces methyl ketones at a bulk production capacity of 1× 107 kg/a for a total cost of 6.0USD/kg.
Reaction models are essential for understanding chemical reactions, but modeling them is a time-demanding process. Automated reaction space exploration techniques, such as ChemTraYzer-TAD, can simplify this process. However, finding transition states (TS) remains a hurdle. TS geometries are crucial for calculating reaction rate constants. Quantum mechanical methods are computationally expensive for TS geometry searches, while reactive molecular mechanics, like ReaxFF, offer faster calculations. Accurate TS searches require second derivatives of energy. State-of-the-art ReaxFF implementations can provide these derivatives only through finite differentiation (FD), which introduces noise. Automatic differentiation (AD) can provide more accurate second derivatives. Hence, this work integrates AD into the classical molecular dynamics code LAMMPS for calculations of second derivatives, presenting ADfied LAMMPS. By interfacing ADfied LAMMPS with the Gaussian computational chemistry suite, LMP-Gau is developed, enabling efficient geometry optimization, frequency calculations, and reaction path following for any force field. LMP-Gau demonstrates improved energy minimization for stable molecules in comparison to standard LAMMPS methods. It is also used successfully to find transition states in 1,3-dioxolane oxidation, demonstrating improved convergence with AD compared to FD.
Ab initio molecular dynamics (AIMD)-derived vibrational spectra provide a promising route toward calibration-free quantitative spectroscopy when combined with indirect hard modeling (IHM). Reliability of AIMD-based spectra for bulk-phase can be compromised by incomplete sampling, electronic-structure errors, and frequency shifts relative to experiment. In this work, an improved AIMD-IHM framework is presented that addresses these limitations through benchmarking comparison of the BLYP and B3LYP/ADMM functional methods to assess their relative accuracy, cluster-resolved sampling analysis based on hydrogen-bond kinetics, and a gas-phase-anchored vibrational frequency scaling strategy. The methodology is demonstrated for aqueous acetic acid, a strongly hydrogen-bonded system characterized by transient molecular associations and proton-sharing motifs. Raman spectra generated from bulk-phase AIMD simulations using BLYP and B3LYP/ADMM are benchmarked against experiment, revealing that the computationally efficient BLYP functional outperforms B3LYP/ADMM in reproducing experimental vibrational frequencies, with a root mean square error (RMSE) of 91 cm-1 compared to 155 cm-1 for the volumetric fraction of 0.2. A region-specific scaling procedure derived from gas-phase data significantly reduces the RMSE of BLYP-based bulk-phase spectra. Hydrogen-bond cluster analysis via reactive-flux indicates that the AIMD trajectories are sufficiently sampled for all relevant molecular motifs and provide quantitative evidence that the spectra are unlikely to be biased by undersampling. When integrated into IHM, the scaled AIMD-derived spectra determine experimental mixture compositions with an RMSE of 0.021 without any experimental calibration, representing an improvement over unscaled spectra, which showed an RMSE of 0.034. The proposed framework enhances the predictive accuracy of AIMD-IHM and extends its applicability to strongly interacting liquid systems.
Accurate prediction of toluene/water partition coefficients of neutral species is crucial in drug discovery and separation processes; however, data-driven modeling of these coefficients remains challenging due to limited available experimental data. To address the limitation of available data, we apply multi-fidelity learning approaches leveraging a quantum chemical dataset (low fidelity) of approximately 9000 entries generated by COSMO-RS and an experimental dataset (high fidelity) of about 250 entries collected from the literature. We explore the transfer learning, feature-augmented learning, and multi-target learning approaches in combination with graph neural networks, validating them on two external datasets: one with molecules similar to training data (EXT-Zamora) and one with more challenging molecules (EXT-SAMPL9). Our results show that multi-target learning significantly improves predictive accuracy, achieving a Root-Mean-Square Error (RMSE) of 0.44 logP units for the EXT-Zamora, compared to an RMSE of 0.63 logP units for single-task models. For the EXT-SAMPL9 dataset, multi-target learning achieves an RMSE of 1.02 logP units, indicating reasonable performance even for more complex molecular structures. These findings highlight the potential of multi-fidelity learning approaches that leverage quantum chemical data to improve toluene/water partition coefficient predictions and address challenges posed by limited experimental data. We expect applicability of the methods used beyond just toluene/water partition coefficients.
Predicting the physicochemical properties of ionizable solutes, including solubility and lipophilicity, is of broad significance. Such predictions rely on the accurate determination of solvation free energies for ions. However, the limited availability of high-quality reference data poses a challenge in developing accurate, inexpensive computational prediction methods. In this study, we address both issues of data quality and availability. We present three databases and models related to ionic phenomena: (1) 8,241 pKa data points across 8 solvents, (2) 5,536 gas-phase acidities from DLPNO-CCSD(T) QM calculations, and (3) 6,090 solvation free energies of anions across 8 solvents obtained from a thermodynamic cycle. We also report 6,088 solvation free energies of neutral conjugate solutes computed using the COSMO-RS method. The pKa data were obtained from the iBonD database, cleaned, and combined with a separate compilation of trustworthy reference pKa data. Gas-phase acidities were computed for most of the acids present in the pKa corpus. Leveraging these data, we compiled values for solvation free energies of anions. We then trained several graph neural network models, which can be used as an alternative to QM approaches to quickly estimate these properties. The pKa and gas-phase acidity models accept reaction SMILES strings of the acid dissociation as inputs, whereas the solvation energy model accepts the SMILES string of the anion. Our microscopic pKa model achieves good accuracy, with an overall test mean average error of 0.58 units on unseen solutes and 0.59 on the SAMPL7 challenge (the lowest error so far among multisolvent models). Our gas-phase acidity model had mean absolute errors slightly above 2 kcal mol-1 when evaluated against experimental data. The anionic solvation free energy model had mean absolute errors of less than 3 kcal mol-1 in several test evaluations, comparable to (though less reliable than) several widely used QM-based solvation models. The models and data are free and publicly available at doi.org/10.5281/zenodo.13987781.
Liquid organic hydrogen carriers (LOHCs) can store and transport hydrogen by chemical bonding. Benzyltoluene (H0-BT) is an attractive LOHC that can take up 12 H per carrier molecule. The chemical equilibrium favors hydrogenation at lower temperatures and higher pressures. In this work, we study hydrogenation kinetics at 125-200 degrees C and 0.3-30 bar H2. We perform ab initio calculations of all isomers of H0-BT and its (partially) hydrogenated forms to compute chemical equilibrium compositions. Despite hydrogenation being exothermic, full hydrogenation is thermodynamically possible almost up to the boiling temperature of the LOHC. Based on the obtained results, the tradeoff between the degree of hydrogenation and reducing the operating temperature is discussed.
Developing chemical combustion mechanisms for novel bio- and e-fuels is a challenge, given that many established mechanism analogies originate from fossil fuels. This work studies the ability of two new reaction network exploration methods, ChemTraYzer-TAD and PESmapping, to suggest potentially important reaction paths to combustion mechanisms, highlighting the differences and similarities of the two methods. In the reaction space of the ethyl-2-yl formate radical combustion, all expected important reactions are found, with many additional reactions suggested by one or the other method. This shows that the combination of both methods provides an optimal exploration result, i.e. both expected and novel reaction pathways.
In this paper, we present an alternative method to ab-initio molecular dynamics (AIMD) simulations for Raman spectra calculations of water molecules which can be computationally expensive. We offer a more efficient method for spectra calculation by utilizing neural network potential (NNP) to reduce computational costs while maintaining accuracy comparable to AIMD simulations. The Deep Polar (DeepPol) model, trained using data from density functional theory simulations, predicts polarizabilities without relying on central atom assignments, allowing for environment-dependent contributions from all atoms. We validate the simulated spectra by comparing results to both AIMD simulations and experimental Raman spectra, analyzing the temperature dependence of the OH stretching band. Key parameters such as sampling time, correlation depth, and system size are systematically investigated to understand their effects on spectral outcomes. The findings demonstrate that machine learning potentials, when integrated with molecular dynamics simulations, provide a computationally efficient framework for simulating Raman spectra, with potential applications beyond water systems.
In this study, we introduce openCOSMO-RS 24a, an improved version of the open-source COSMO-RS model parameterized using quantum chemical calculations from ORCA 6.0, leveraging a comprehensive dataset that includes solvation free energies, partition coefficients, and infinite dilution activity coefficients for various solutes and solvents mainly at 25 degrees C. This is the first version of the model also capable of predicting solvation free energies based on ORCA calculations. Additionally, we develop a Quantitative Structure-Property Relationships model to predict molar volumes of the solvents, an essential requirement for predicting solvation free energies and partition coefficients from structure alone. Our results show that openCOSMO-RS 24a achieves an average absolute deviation of 0.45 kcal mol1 for solvation free energies, 0.76 for the logarithm of the partition coefficients, and 0.51 for the logarithm of infinite dilution activity coefficients, demonstrating improvements over the previous openCOSMO-RS 22 parameterization and comparable results to COSMOtherm 24 BP-TZVP. The user interface was extended to be able to use it as solvation model directly from within ORCA 6.0 or from the command line to provide researchers with a robust tool for applications in chemical and materials science.
Polyelectrolytes often display good solubility in water but not in organic solvents, a feature that limits their applications in nonaqueous media such as hand sanitizers. Here, we show that this limitation can be overcome by tuning the counterion-solvent affinity. To this end, the solubility and chain conformation of carboxymethylcellulose (CMC) salts with different organic counterions in a variety of solvents were studied by employing the Hansen solubility parameter (HSP) framework and small-angle X-ray scattering (SAXS), respectively. The solubility phase mapping demonstrates an increase in the soluble region in HSP space for the polyelectrolyte to encompass more solvents as the counterion side arm length increases or if the side arm is substituted with a large functional group, while substituting the central atom does not change the solubility, suggesting that the solubility is mainly influenced by the interaction between the peripheral atoms and the solvents. Scattering measurements revealed that for a given solvent, the nature of the counterion does not influence the conformation of chains in solution, as seen by the independence of the stretching parameter B on counterion type.
More sustainable chemical processes require the selection of suitable molecules, which can be supported by computer-aided molecular design (CAMD). CAMD often generates and evaluates molecular structures using genetic algorithms. However, genetic algorithms can suffer from slow convergence, and might yield suboptimal solutions. In response to these challenges, this work presents a method to fine-tune a genetic algorithm for CAMD. The proposed method builds on the COSMO-CAMD framework that utilizes a genetic algorithm for solving optimization-based molecular design problems and COSMO-RS for predicting physical properties of molecules. The key idea of the proposed method is to integrate results from a fast large-scale molecular screening into the molecular design framework through an automated fragmentation procedure. By generating a promising initial population and constructing a tailored fragment library, our method enables a targeted initialization of the genetic algorithm, referred to as warm-start. The proposed method is applied in two case studies to design solvents for extracting γ-valerolactone and phenol, respectively, from aqueous solutions. Compared to the benchmark method, the warm-started COSMO-CAMD framework achieves a 70% faster convergence, discovers 4-fold more top-performing candidate molecules, and identifies seven tailored molecular fragments, culminating in the discovery of two novel solvents specifically for the phenol case. The optimal solvent is found in all computational runs. Overall, the warm-started COSMO-CAMD framework significantly improves efficiency, effectiveness, and robustness of molecular design.
Charged microgels are promising for different applications due to their multiresponsive nature. Achieving tailored microgel properties necessitates synthesis monitoring and computational modeling. However, process analytics of multicomponent solutions occurring during charged microgel synthesis face challenges. Furthermore, existing modeling approaches employ steady-state models while lacking fully determined kinetic parameter values. In this contribution, we use Raman spectroscopy and indirect hard modeling to measure concentrations of samples obtained during the synthesis of charged microgels and predict the monomer content. Additionally, we derive a dynamic synthesis model for charged microgels including pH dependence. Based on the data obtained from Raman spectroscopy, calorimetry, and quantum mechanical computations, we estimate a complete set of reaction parameter values for the synthesis of poly(N-isopropylacrylamide-co-methacrylic) acid microgels. The dynamic model allows simulation of the microgel synthesis, revealing effects of pH changes. Thus, this work represents an important step toward the model-based production of tailored multiresponsive microgels.
Biotechnologically produced methyl ketones can be a sustainable, safe, and less toxic biofuel candidate with efficient and clean combustion properties and compatibility with the fuel infrastructure.
The Perturbed Chain Polar Statistical Associating Fluid Theory (PCP-SAFT) equation of state (EoS) is widely used to predict fluid-phase thermodynamics, but parameterization of PCP-SAFT for individual molecules is often challenging. We propose a machine learning framework called ML-SAFT that can turn experimental data in predictive models of PCP-SAFT parameters. We demonstrate methods for automated large scale regression of PCP-SAFT parameters and thus create a large PCP-SAFT parameter dataset in the literature. We then evaluate several machine learning architectures for predicting PCP-SAFT parameters. We find that our best model provides accurate predictions for a wider range of molecules than existing predictive methods with 40 % average absolute deviation (% AAD) in vapor pressure predictions and 8 % AAD in density predictions.
Accurate thermochemistry computations often require proper treatment of torsional modes. The one-dimensional hindered rotor model has proven to be a computationally efficient solution, given a sufficiently accurate potential energy surface. Methods that provide potential energies at various compromises of uncertainty and computational time demand can be optimally combined within a multifidelity treatment. In this study, we demonstrate how multifidelity modeling leads to (1) smooth interpolation along low-fidelity scan points with uncertainty estimates, (2) inclusion of high-fidelity data that change the energetic order of conformations, and (3) predicting best next-point calculations to extend an initial coarse grid. Our diverse application set comprises molecules, clusters, and transition states of alcohols, ethers, and rings. We discuss limitations for cases in which the low-fidelity computation is highly unreliable. Different features of the potential energy curve affect different quantities. To obtain "optimal" fits, we apply strategies ranging from simple minimization of deviations to developing an acquisition function tailored for statistical thermodynamics. Bayesian prediction of best next calculations can save a substantial amount of computation time for one- and multidimensional hindered rotors.
Liquid phase oxidation of furfural using hydrogen peroxide offers a promising route for bio-based C4 furanones and diacids; however, only dilute water-based process designs have been previously suggested that have limited techno-economic potential. In this study, a conceptual process design is presented, where aqueous furfural is extracted using an organic solvent, coupled with peroxide oxidation and product recovery in the presence of the solvent. To address the problem of solvent selection, the COSMO-RS-based solvent screening framework is applied, where quantum mechanics-based thermodynamics are utilized in pinch-based process models. About 2500 solvent candidates were identified as feasible. Focusing on a set of 400 solvent candidates revealed energy consumption values (Qreb,tot/m(center dot)prod recov) between approximately 2 MWh/tonne and 33 MWh/tonne, signifying the potential of the solvent-based process in outperforming the reference aqueous process (49.4 MWh/tonne). The study provides potential solvent candidates and future directions to consider in more costly computational and experimental efforts.
Creation of complex chemical mechanisms for hydrocarbon pyrolysis and combustion is challenging due to the large number of species and reactions involved. Reactive molecular dynamics (RMD) enables the simulation of thousands of reactions and the discovery of previously unknown components of the reaction network. However, due to the inherent imprecision of reactive force fields, it is necessary to verify RMD-obtained reaction paths using more accurate methods such as Density Functional Theory (DFT). We demonstrate a method for identification and confirmation of reaction pathways from RMD that supplement an established mechanism, using the example of benzene formation from n-heptane and iso-octane pyrolysis. We establish a validation workflow to extract reaction geometries from RMD and optimize transition states using the Nudged-Elastic-Band method on semi-empirical and quantum mechanical levels of theory. Our findings demonstrate that the widely recognized ReaxFF parameterization, CHO2016, can identify known pathways from a established soot formation mechanism while also indicating new ones. We also show that CHO2016 underestimates hydrogen migration barriers by up to 40kcalmol-1$40\,{\rm {kcal\,mol}}<^>{-1}$ as compared to DFT and can lower activation barriers significantly for spin-forbidden reactions. This highlights the necessity for validation or potentially even reparametrization of CHO2016.
Biohybrid fuels are a promising solution for making the transportation sector more environmentally friendly. One such interesting fuel candidate is 1,3-dioxolane, which can be produced from inedible biomass. However, very little kinetics data are available for the low-temperature oxidation of this fuel molecule. To remedy this, we present the reaction kinetics of O2 addition to 1,3-dioxolanyl radicals in this work. All energies have been calculated at the DLPNO-CCSD(T)/CBS//B2PLYPD3BJ/6-311+g(d,p) level of theory. Temperature- and pressure-dependent reaction rate constants have been calculated with the RRKM/master equation method. The effects of heterocyclic oxygen atoms and ring strain on the low-temperature oxidation of 1,3-dioxolane are also compared to that of similar fuel molecules containing five heavy atoms: cyclopentane, tetrahydrofuran, and diethyl ether (DEE). The ring-opening β-scission reactions of the dioxolane hydroperoxy species are found to be the most dominant pathways following the oxidation of 1,3-dioxolanyl radicals. The heterocyclic oxygen atoms in 1,3-dioxolane weaken its C-O bonds, which leads to low barrier heights of the ring-opening reactions. Ring strain in 1,3-dioxolane increases the barriers for isomerization reactions of peroxy radicals compared to the similar reactions of DEE, which has a chain structure.