X-ray absorption near edge structure (XANES) spectroscopy is a powerful method to probe the oxidation state and local structure of metals in catalytic materials. However, it suffers from the lack of unbiased data analysis protocols. Machine learning (ML) overcomes human-related factors by uncovering relevant spectrum-structure relationships and subsequent cross-validation analysis. The bottlenecks in the automatic processing of experimental data are the lack of chemically diverse XANES reference libraries and systematic differences between theory and experiment. Therefore, compiling experimental reference libraries across the periodic table and rational application of ML methodology to small (in terms of data science) training datasets becomes increasingly important. This work revises the classical XANES fingerprint analysis by database augmentation, feature extraction, cross-validation, and uncertainty analysis. We apply the developed methodology to decipher the oxidation state and local coordination of supported vanadium-oxo species (VOx), which change their structure participating in oxidative dehydrogenation catalysis. The developed library and instruments for analysis may serve as a starting point for a unified platform of fingerprint XANES data analysis.
Metal nanoparticles are widely used as heterogeneous catalysts to activate adsorbed molecules and reduce the energy barrier of the reaction. Reaction product yield depends on the interplay between elementary processes: adsorption, activation, desorption, and reaction. These processes, in turn, depend on the inlet gas composition, temperature, and pressure. At a steady state, the active surface sites may be inaccessible due to adsorbed reagents. Periodic regime may thus improve the yield, but the appropriate period and waveform are not known in advance. Dynamic control should account for surface and atmospheric modifications and adjust reaction parameters according to the current state of the system and its history. In this work, we applied a reinforcement learning algorithm to control CO oxidation on a palladium catalyst. The policy gradient algorithm was trained in the theoretical environment, parametrized from experimental data. The algorithm learned to maximize the CO2 formation rate based on CO and O2 partial pressures for several successive time steps. Within a unified approach, we found optimal stationary, periodic, and nonperiodic regimes for different problem formulations and gained insight into why the dynamic regime can be preferential. In general, this work contributes to the task of popularizing the reinforcement learning approach in the field of catalytic science.
Ti-based molecules and materials are ubiquitous, and play a major role in both homogeneous and heterogeneous catalytic processes. Understanding the electronic structures of their active sites (oxidation state, local symmetry and ligand environment) is key to developing molecular-level structure-property relationships. In that context, X-ray absorption spectroscopy (XAS) offers a unique combination of element selectivity and sensitivity to local symmetry. Commonly, for early transition metals such as Ti, K-edge XAS is applied for in situ characterization and subsequent structural analysis with high sensitivity towards tetrahedral species. Ti L2,3-edge spectroscopy is in principle complementary and offers specific opportunities to interrogate the electronic structure of five-and six-coordinated species. It is, however, much more rarely implemented, because the use of soft X-rays implies ultra-high vacuum conditions. Furthermore, the interpretation of the data can be challenging. Here, we show how Ti L2,3-edge spectroscopy can help to obtain unique information about both homogenous and heterogeneous epoxidation catalysts and to develop a molecular-level relationship between spectroscopic signatures and electronic structures. Towards this goal, we first establish a spectral library of molecular Ti reference compounds, comprising various coordination environments with mono- and dimeric Ti species having O, N and Cl-ligands. We next implemented a computational methodology based on multiplet ligand field theory and maximally localized Wannier orbitals benchmarked on our library to understand Ti L2,3-edge spectroscopic signatures. We finally used this approach to track and predict spectra of catalytically relevant intermediates, focusing on Ti-based olefin epoxidation catalysts.
X-ray absorption spectroscopy (XAS) has been central to the study of the Phillips polymerization catalyst (CrO3/SiO2).
The ethylene polymerization Phillips catalyst has been employed for decades and is central to the polymer industry. While Cr(III) alkyl species are proposed to be the propagating sites, there is so far no direct experimental evidence for such proposal. In this work, by coupling Surface organometallic chemistry (SOMC), EPR spectroscopy, and machine learning-supported XAS studies, we have studied the electronic structure of well-defined silica-supported Cr(III) alkyls, and identified the presence of several surface species from high to low spin Cr(III), associated with different coordination environments. Notably, low-spin Cr(III) sites are shown to participate in ethylene polymerization, indicating that similar Cr(III) alkyl species could be involved in the related Phillips catalyst.
Application of machine learning (ML) algorithms to spectroscopic data has a great potential for obtaining hidden correlations between structural information and spectral features. Here, we apply ML algorithms to theoretically simulated infrared (IR) spectra to establish the structure-spectrum correlations in zeolites. Two hundred thirty different types of zeolite frameworks were considered in the study whose theoretical IR spectra were used as the training ML set. A classification problem was solved to predict the presence or absence of possible tilings and secondary building units (SBUs). Several natural tilings and SBUs were also predicted with an accuracy above 89%. The set of continuous descriptors was also suggested, and the regression problem was also solved using the ExtraTrees algorithm. For the latter problem, additional IR spectra were computed for the structures with artificially modified cell parameters, expanding the database to 470 different spectra of zeolites. The resulting prediction quality above or close to 90% was obtained for the average Si-O distances, Si-O-Si angles, and volume of TO4 tetrahedra. The obtained results provide new possibilities for utilization of infrared spectra as a quantitative tool for characterization of zeolites.
X-ray absorption spectroscopy (XAS) is one of the most powerful characterization techniques, that has been intensively employed to study the Phillips polymerization catalyst (CrO3/SiO2). While Cr K-edge XAS signatures are used to evaluate the nature of surface (active) sites, they are highly sensitive to oxidation state, geometry and types of ligands, making interpretation challenging. In the specific case of CrO3/SiO2, CO has been particularly used both as a reductant to generate the expected low valent Cr sites and a probe to understand surface Cr sites. Considering the electronic properties of CO, a strong sigma-donor and pi-acceptor ligand, one may wonder the impact of the coordination of CO on Cr on its XAS signature. We herein built a molecular low-valent Cr library bearing isocyanide ligands, which mimic CO as its isoelectronic counterpart, as a model of low-valent Cr sites interacting with pi-acceptor ligand. Cr K-edge XAS augmented with DFT calculations elucidated the profound effect of isocyanide ligand on both XANES and EXAFS regions giving a rise to characteristic features as well as the significant stabilization of low-spin Cr(II/III) species, which potentially alter the ease of interpretation of XAS spectra. Taking the herein demonstrated effect of pi-acceptor ligand into account, experimental Cr K-edge spectra of reduced Phillips catalyst at different temperatures, with/without interaction with CO, were nicely reproduced.
Gold nanoparticles represent an important class of functional nanomaterials for optoelectronics, biomedical applications, and catalysis. Therefore, controllable synthesis of nanoparticles with specified size and shape is important. Though reduction of gold ions is quite a simple process and may be performed with many different protocols, the reproducibility of the results and transfer of protocols between independent research groups remains a challenging task. Machine learning analysis based on statistical approaches is hardly applicable to the published data, since most of the researchers report only successful syntheses. In this work, we apply uniform sampling of the reaction parameter space. The concentrations of gold precursor, reducing agent, and surfactant were varied via an improved Latin hypercube sampling, and each run was performed under in situ UV-vis control. Based on the resulting set of optical spectra, we address the relevant chemical questions about nanoparticle formation, their shape, and period of growth. Our work demonstrates a data driven approach applied to the space of reaction parameters in a limited available set of experiments.
CuCrP 2 S 6 , a van der Waals magnet having stacked layers of 2D honeycomb lattice made of CuS 3 triangles and CrS 6 octahedra, exhibits an A‐type antiferromagnetic order with the Néel temperature ( T N ) = 32 K. Upon in‐plane magnetic field ( H ) being applied below T N , H ‐induced modulation of the c *‐axis electric polarization (Δ P c* ) is found at fields lower than the saturation field µ 0 H S = 6.1 T, at which a forced ferromagnetic alignment sets in. Based on the symmetry analyses and dependence of Δ P c* on H and the azimuthal angle of applied H direction, a microscopic origin of the magnetoelectric (ME) coupling is attributed to the spin‐direction‐dependent p – d hybridization that is allowed due to the presence of off‐centered Cr 3+ octahedra. A comparative study on CuCrP 2 Se 6 , however, finds no H ‐induced P modulation due to cancellation of P between neighboring layers with the doubling of a crystallographic unit cell at T N . As the p – d hybridization mechanism allows generation of P in a single Cr atom–ligand pair, the results imply that large ME coupling should exist even in a single layer limit of CuCrP 2 S 6 .
Microscopic tissue analysis is the key diagnostic method needed for disease identification and choosing the best treatment regimen. According to the Global Cancer Observatory, approximately two million people are diagnosed with colorectal cancer each year, and an accurate diagnosis requires a significant amount of time and a highly qualified pathologist to decrease the high mortality rate. Recent development of artificial intelligence technologies and scanning microscopy introduced digital pathology into the field of cancer diagnosis by means of the whole-slide image (WSI). In this work, we applied deep learning methods to diagnose six types of colon mucosal lesions using convolutional neural networks (CNNs). As a result, an algorithm for the automatic segmentation of WSIs of colon biopsies was developed, implementing pre-trained, deep convolutional neural networks of the ResNet and EfficientNet architectures. We compared the classical method and one-cycle policy for CNN training and applied both multi-class and multi-label approaches to solve the classification problem. The multi-label approach was superior because some WSI patches may belong to several classes at once or to none of them. Using the standard one-vs-rest approach, we trained multiple binary classifiers. They achieved the receiver operator curve AUC in the range of 0.80–0.96. Other metrics were also calculated, such as accuracy, precision, sensitivity, specificity, negative predictive value, and F1-score. Obtained CNNs can support human pathologists in the diagnostic process and can be extended to other cancers after adding a sufficient amount of labeled data.
Catalytic properties of noble-metal nanoparticles (NPs) are largely determined by their surface morphology. The latter is probed by surface-sensitive spectroscopic techniques in different spectra regions. A fast and precise computational approach enabling the prediction of surface–adsorbate interaction would help the reliable description and interpretation of experimental data. In this work, we applied Machine Learning (ML) algorithms for the task of adsorption-energy approximation for CO on Pd nanoclusters. Due to a high dependency of binding energy from the nature of the adsorbing site and its local coordination, we tested several structural descriptors for the ML algorithm, including mean Pd–C distances, coordination numbers (CN) and generalized coordination numbers (GCN), radial distribution functions (RDF), and angular distribution functions (ADF). To avoid overtraining and to probe the most relevant positions above the metal surface, we utilized the adaptive sampling methodology for guiding the ab initio Density Functional Theory (DFT) calculations. The support vector machines (SVM) and Extra Trees algorithms provided the best approximation quality and mean absolute error in energy prediction up to 0.12 eV. Based on the developed potential, we constructed an energy-surface 3D map for the whole Pd55 nanocluster and extended it to new geometries, Pd79, and Pd85, not implemented in the training sample. The methodology can be easily extended to adsorption energies onto mono- and bimetallic NPs at an affordable computational cost and accuracy.
X-ray absorption near-edge structure (XANES) spectroscopy is a powerful characterization technique that is sensitive to both three-dimensional (3D) geometry and the electronic state of the selected element. In this work, we have suggested a set of structural descriptors that can be used to characterize the state of palladium nanoparticles in hydrogenation reactions and explored the possibility of their extraction from Pd K-edge XANES spectra. A theoretical spectral database was calculated for palladium atoms in the bulk and at the (111) surface with variable Pd-Pd interatomic distances. Carbon and hydrogen atoms randomly occupied octahedral interstitial sites for different H/Pd and C/Pd ratios. The presence of hydrogen and hydrocarbon molecules adsorbed at the surface was also considered. The obtained spectral database was subjected to the principal component analysis (PCA) to estimate the number of strongly contributing components and the multivariate curve resolution (MCR) approach to deconvolve the whole set of data into the XANES spectra of "pure" species and their concentrations. The latter were also used as descriptors of spectra, and machine learning (ML) algorithms were then trained to predict them based on the descriptors of structure and vice versa. We have shown that some of the structural parameters, namely, the concentration of surface-adsorbed molecules, have minor effects on the spectra and cannot be predicted. For interatomic distances, their averaged value can be extracted with good prediction quality based on only one MCR concentration, while independent prediction of the distances in the bulk or at the surface gives unsatisfactory results. Finally, we constructed a new set of structural descriptors that have direct relevance to the MCR components.
Модели многозначной классификации возникают в различных сферах современной жизни, что объясняется всё большим количеством информации, требующей оперативного анализа. Одним из математических методов решения этой задачи является модульный метод, на первом этапе которого для каждого класса строится некоторая ранжирующая функция, упорядочивающая некоторым образом все объекты, а на втором этапе для каждого класса выбирается оптимальное значение порога, объекты с одной стороны которого относят к текущему классу, а с другой — нет. Пороги подбираются так, чтобы максимизировать целевую метрику качества. Алгоритмы, свойства которых изучаются в настоящей статье, посвящены второму этапу модульного подхода — выбору оптимального вектора порогов. Этот этап становится нетривиальным в случае использования в качестве целевой метрики качества $F$-меры от средней точности и полноты, так как она не допускает независимую оптимизацию порога в каждом классе. В задачах экстремальной многозначной классификации число классов может достигать сотен тысяч, поэтому исходная оптимизационная задача сводится к задаче поиска неподвижной точки специальным образом введенного отображения $\boldsymbol V$, определенного на единичном квадрате на плоскости средней точности $P$ и полноты $R$. Используя это отображение, для оптимизации предлагаются два алгоритма: метод линеаризации $F$-меры и метод анализа области определения отображения $\boldsymbol V$. На наборах данных многозначной классификации разного размера и природы исследуются свойства алгоритмов, в частности зависимость погрешности от числа классов, от параметра $F$-меры и от внутренних параметров методов. Обнаружена особенность работы обоих алгоритмов для задач с областью определения отображения $\boldsymbol V$, содержащей протяженные линейные участки границ. В случае когда оптимальная точка расположена в окрестности этих участков, погрешности обоих методов не уменьшаются с увеличением количества классов. При этом метод линеаризации достаточно точно определяет аргумент оптимальной точки, а метод анализа области определения отображения $\boldsymbol V$ — полярный радиус.
Unveiling the nature and the distribution of surface sites in heterogeneous catalysts, and for the Phillips catalyst (CrO3/SiO2) in particular, is still a grand challenge despite more than 60 years of research. Commonly used references in Cr K-edge XANES spectral analysis rely on bulk materials (Cr-foil, Cr2O3) or molecules (CrCl3) that significantly differ from actual surface sites. In this work, we built a library of Cr K-edge XANES spectra for a series of tailored molecular Cr complexes, varying in oxidation state, local coordination environment, and ligand strength. Quantitative analysis of the pre-edge region revealed the origin of the pre-edge shape and intensity distribution. In particular, the characteristic pre-edge splitting observed for Cr(III) and Cr(IV) molecular complexes is directly related to the electronic exchange interactions in the frontier orbitals (spin-up and -down transitions). The series of experimental references was extended by theoretical spectra for potential active site structures and used for training the Extra Trees machine learning algorithm. The most informative features of the spectra (descriptors) were selected for the prediction of Cr oxidation states, mean interatomic distances in the first coordination sphere, and type of ligands. This set of descriptors was applied to uncover the site distribution in the Phillips catalyst at three different stages of the process. The freshly calcined catalyst consists of mainly Cr(VI) sites. The CO-exposed catalyst contains mainly Cr(II) silicates with a minor fraction of Cr(III) sites. The Phillips catalyst exposed to ethylene contains mainly highly coordinated Cr(III) silicates along with unreduced Cr(VI) sites.
In this work, we propose a new method for the analysis of time-resolved X-ray absorption near edge structure (XANES) spectra. It allows to decompose an experimental dataset as the product of two matrices: a pure spectral matrix, composed by XANES spectra associable to well-defined chemical species/sites, and their related concentration profiles. This method combines the principal component analysis and the application of a transformation matrix whose elements are directly accessible by the user. We demonstrate the potential of this approach applying it to a series of XANES spectra acquired during the direct conversion of methane to methanol (DMTM) over a Cu-exchanged zeolite characterized by the ferrierite topology. Possibilities and limitations of this methodology are discussed together with a critical comparison with the Multivariate Curve Resolution Alternating Least Squares (MCR-ALS) algorithm that, in the field of X-ray absorption spectroscopy (XAS), is imposing itself as a widely used method for spectral decomposition.
Nowadays a multi-label classification problem arises in different areas for which the significant amount of data has been gained. This problem can be viewed as the one comprising two steps: training some ranking function sorting instances in each class and defining the optimal number of predictions for it. This paper is devoted to the second step of the optimal threshold selection while maximizing the F-macro measure. To do so, we reduce the multi-dimensional problem to the two-dimensional problem of finding a fixed point of a specifically introduced transformation defined on a unit square. We suggest the algorithm of finding the vector of optimal thresholds based on the domain analysis of the introduced transformation. Moreover, we provide the complexity estimations of the proposed algorithm. We evaluate the algorithm on the extreme classification benchmark WikiLSHTC-325K comparing its performance with some baseline results.
Element selectivity and possibilities for in situ and operando applications make X-ray absorption spectroscopy a powerful tool for structural characterization of catalysts. While determination of coordination numbers and interatomic distances from extended spectral region is rather straightforward, analysis of X-ray absorption near-edge structure (XANES) spectra remains a highly debated and topical problem. The latter region of spectra is shaped depending on the local 3D geometry and electronic structure. However, there is no straightforward procedure for the unambiguous extraction of these parameters. This work gives a critical vision on the amount of information that can be practically extracted from Pd K-edge XANES spectra measured under in situ and operando conditions, in which adsorption of reactive molecules at the surface of palladium with further formation of subsurface and bulk palladium carbides are expected. We investigate how particle size, concentration of carbon impurities, and their distribution in the bulk and at the surface of palladium particles affect Pd K-edge XANES features and to which extend they should be implemented in the theoretical model to adequately reproduce experimental data. Then, we show how the step-by-step increasing the complexity of the theoretical model improves the agreement with experiment. Finally, we suggest a set of formal descriptors relevant to possible structural diversity and construct a library of theoretical spectra for machine-learning-based analysis of the experimental data.
Formation of gold nanoparticles (NPs) from the mixture containing NaAuCl4 as a precursor, octadecene as a solvent and oleylamine as a reducing agent was studied in situ by means of optical and X-ray spectroscopies. Dynamic light scattering (DLS) revealed the presence of initial aggregates of 500 nm average size which split into nanoparticles of about 8 nm width shortly after the reduction from Au3+ to Au+ has been completed. Based on Density Functional Theory (DFT) simulations and analysis of X-ray absorption spectra (XAS) we identified the structure of Au3+ and Au+ gold complexes. Quantitative analysis shows that Au NPs formation proceeds in following steps: substitution of chlorine ligands in Au3+ complex by four oleylamines; reduction of Au3+ to Au+ coordinated by two oleylamines. Latter process occurs in oleylamine micelles in octadecene. The third step is a fragmentation of large micelles into smaller ones shortly after reduction Au3+ to Au+, and subsequent slow growth of Au NPs via reduction of Au+ to Au0.
•Coordinate-wise maximum for the analyzed measure may not be maximum in usual sense.•The proposed fixed point method can localize all maximums of the measure.•The method works the more precise, the bigger class count is.•The approach is successfully applied to the real-world datasets.•The difference from similar macro F measure is investigated.