Prediction of the elution order of close structural analogs and isomers is a critical step in plant metabolites dereplication. The application of machine learning (ML) is an efficient approach to automate peak annotation by implementing structure-retention relationships. In this study, four ML models were trained to predict the elution order flavonoid derivative pairs containing a flavone aglycone bearing hydroxy and methoxy substituents. Additionally, retention times were indirectly estimated from model scores using linear regression and linear interpolation. Two data sets, comprising 51 compounds (1275 pairs) and 48 compounds (356 pairs), were constructed from available literature data. These data sets were used to explore alternative model training strategies and to conduct internal and external validation. A specially designed molecular fingerprint was employed to encode structural features of the flavone scaffold and its substituents, optimizing the representation for this widespread class of phytochemicals, taken as an example. Ranking neural network-based (NN) models with binary cross-entropy (BCE) and margin ranking (MR) loss functions employed a simplified version of the fingerprint, whereas logistic regression models were tested using both a condensed (20-bit) and an extended (92-bit) fingerprint that incorporated interaction between substituents. The pairwise error rates for elution order prediction were predominantly below 10%, demonstrating reliable performance under reversed-phase LC conditions with an acetonitrile gradient. Linear regression slightly outperformed the other models with the statistical significance indicated by Friedman and Wilcoxon tests. Although overall performance metrics were comparable, the use of a large uniform data set was found to be preferable over fragmented literature-derived data. The positive and negative effects of hydroxy and methoxy groups in different positions, along with their interactions on chromatographic retention, were analyzed following the visualization of model weights.
The structural study of plant viruses is of great importance to reduce the damage caused by these agricultural pathogens and to support their biotechnological applications. Nowadays, X-ray crystallography, NMR spectroscopy and cryo-electron microscopy are well accepted methods to obtain the 3D protein structure with the best resolution. However, for large and complex supramolecular structures such as plant viruses, especially flexible filamentous ones, there are a number of technical limitations to resolving their native structure in solution. In addition, they do not allow us to obtain structural information about dynamics and interactions with physiological partners. For these purposes, small-angle X-ray scattering (SAXS) and atomic force microscopy (AFM) are well established. In this review, we have outlined the main principles of these two methods and demonstrated their advantages for structural studies of plant viruses of different shapes with relatively high spatial resolution. In addition, we have demonstrated the ability of AFM to obtain information on the mechanical properties of the virus particles that are inaccessible to other experimental techniques. We believe that these under-appreciated approaches, especially when used in combination, are valuable tools for studying a wide variety of helical plant viruses, many of which cannot be resolved by classical structural methods.
A set of covalently-bonded poly(styrene-divinylbenzene)-based stationary phases with linear polyelectrolyte layers was obtained by one-step epoxy-amine polymerization and examined. The influence of synthesis conditions, such as temperature, duration of the synthesis and quantity of reagents, on ion-exchange capacity and selectivity toward inorganic anions and organic acids was investigated using hydroxide eluent in suppressed ion chromatography (IC) mode. Obtained stationary phases packed in 10-cm long columns allowed the separation of up to 21 anions in 45 min, including mono-, di-, and trivalent organic acids, as well as inorganic anions. For the first time the possibility of using PS-DVB-based polyelectrolyte-grafted stationary phase for separation of watersoluble vitamins, sugars, nucleosides and nitrogenous bases in hydrophilic interaction liquid chromatography (HILIC) mode and alkylbenzenes in reversed phase high performance liquid chromatography (RP HPLC) mode was demonstrated. The obtained phases showed beneficial performance as compared to previously reported mixed-mode adsorbents based on the same substrate.
An Erratum to this paper has been published: https://doi.org/10.1134/S0036024424120021
In the pharmaceutical industry, the need for analytical standards is a bottleneck for comprehensive evaluation and quality control of intermediate and end products. These are complex mixtures containing structurally related molecules. In this regard, chromatographic peak annotation, especially for critical pairs of isomers and closest structural analogs, can be supported by using a Quantitative Structure Retention Relationship (QSRR) approach. In our study, we investigated the fundamental basis of the reversed-phase (RP) retention mechanism for 1141 isomeric compounds from the METLIN SMRT dataset. Nine different descriptor calculation tools combined with different feature selection methods (genetic algorithm (GA), stepwise, Boruta) and machine learning (ML) approaches (support vector machine (SVM), multiple linear regression (MLR), random forest (RF), XGBoost) were applied to provide a reliable molecular structure-based interpretation of RP retention behaviour of the isomeric compounds. Strict internal and external validation metrics were used to select models with the best predictive capabilities (rtest > 0.73, order of elution > 60 %). For the developed models, mean absolute errors were in the range of 60 to 110 s. Stepwise and GA showed the most suitable performance as descriptor selection methods, while SVM and XGBoost modeling gave satisfactory predictive characteristics in most cases. Validation performed on the published experimental data for structurally related pharmaceutical compounds confirmed the best accuracy of MLR modeling in combination with GA feature selection of general physico-chemical properties. The resulting models will be useful for the prediction of separation and identification of structurally related compounds in pharmaceutical analysis, providing a simultaneous understanding of the interaction mechanisms leading to their retention under RP conditions.
New adsorbents based on silica and polystyrene–divinylbenzene (PS–DVB) for hydrophilic interaction liquid chromatography (HILIC) with eremomycin in functional layers were obtained. The chromatographic properties of the new adsorbents were assessed using the Tanaka test for hydrophilic stationary phases and by studying the retention of substances of various classes in HILIC, chiral, and reversed-phase chromatography modes. It was shown that the use of eremomycin to create functional layers leads to an increase in the hydrophilicity of the adsorbents on different types of substrates and ensures the shielding of their charge. Eleven nitrogenous bases, nucleosides with an efficiency of up to 25 000 tp/m, or seven vitamins with an efficiency of up to 40 000 tp/m can be separated on a modified sorbent based on aminopropyl silica, and three different HPLC modes can be implemented on the sorbent with eremomycin based on PS–DVB.
The efficiency of the extraction of the terpene fraction from the greenery of common juniper (Juniperus communis L.) by sub- and supercritical extraction methods is compared with the efficiency of conventional methods. The effect of the extraction parameters on the qualitative composition of the extracts is studied. The components of the obtained extracts are identified by gas chromatography–mass spectrometry. It is determined that the studied extracts are similar in qualitative composition, but differ markedly in the quantitative ratio of groups of components. It is shown that monoterpenes are best extracted with butane under subcritical conditions and by hydrodistillation.
Adsorbents based on various substrates—silica and a copolymer of styrene with divinylbenzene—are developed for the determination of amino acids by hydrophilic interaction liquid chromatography—mass spectrometry. The optimal version of the structure of the functional layer in two series of the obtained stationary phases was chosen, which provides the best hydrophilization for each substrate. Retention mechanisms were studied and the conditions for the mass-spectrometric detection, separation, and determination of 16 amino acids were chosen. The applicability of the obtained adsorbents and a method for determining amino acids for the analysis of soil extracts were estimated.
The experimental design methodology based on central composite design of experiments was applied to compare the retention mechanisms in supercritical fluid chromatography (SFC) and non-aqueous hydrophilic interaction liquid chromatography (NA-HILIC). The selected set consists of 26 compounds that belong to imidazoline and serotonin receptor ligands. The different chemometric tools (multiple linear regression, principal component analysis, parallel factor analysis) were used to examine the retention, as well as to identify the most significant retention mechanisms. The retention mechanism was investigated on two different stationary phases (diol, and mixed-mode diol). In NA-HILIC, the mobile phase contains acetonitrile as a main component, and methanolic solution of ammonium formate (+ 0.1% of formic acid) as a modifier. The same mobile phase modifier was used in SFC, with a difference in the main component of the mobile phase which was CO2. The retention behaviour differs significantly between HILIC and SFC conditions. The retention pattern in HILIC mode was more partition-like, while in SFC the solute-sorbent interactions allowed retention. The retention mechanism between mixed-mode diol and the diol phases varies depending on the applied chromatographic mode, e.g., in HILIC the type of stationary phase significantly affects the elution order, while in SFC this was not the case. The HILIC retention behaviour was influenced by the number of tertiary amines-aliphatic, and N atom-centred fragments in tested compounds. On the other hand, the number of pyrrole and pyridine rings in the structure of the compound showed correlation with their SFC retention, simultaneously increasing the molecular weight and rapid elution of larger compounds. It was found that temperature surprisingly plays a major role in SFC mode. The increase in temperature reduces the relative contribution of enthalpy factors to total retention, so the separation in SFC was more entropy-controlled. For further pharmaceutical research and optimization, the SFC would be considered more beneficial compared to HILIC since it gives good selectivity in separation of chosen impurities.
Представлены результаты изучения эффективности выделения терпеновой фракции из древесной зелени можжевельника обыкновенного (Juniperus Communis L.) методами суб- и сверхкритической экстракции в сопоставлении с традиционными методами. Исследовано влияние параметров проведения экстракции на качественный состав экстрактов. Идентификация компонентов полученных экстрактов проведена методом хромато-масс-спектрометрии. Установлено, что качественные составы исследованных экстрактов имеют сходство, но заметно различаются количественным соотношением групп компонентов. Показано, что наибольшему извлечению монотерпенов способствуют экстракция бутаном в субкритических условиях и гидродистилляция. The paper presents the results of studying the efficiency of the extraction of the terpene fraction from the tree greenery of common juniper (Juniperus Communis L.) by sub-and supercritical extraction methods in comparison with traditional methods. The influence of extraction parameters on the qualitative composition of extracts is studied. Identification of the components of the obtained extracts was carried out by chromatography mass spectrometry. It was found that the qualitative compositions of the studied extracts are similar, but they differ significantly in the ratio of the groups of components. It is shown that the greatest extraction of monoterpenes is facilitated by butane extraction under subcritical conditions and hydrodistillation.
Plant samples are potential sources of physiologically active secondary metabolites and their classification is an extremely important task in traditional medicine and other fields of research. In the production of herbal drugs, different plant parts of the same or related species can serve as adulterants for primary plant material. The use of highly informative and relatively easily accessible tools, such as liquid chromatography and low-resolution mass spectrometry, helps to solve these tasks by means of fingerprint analysis. In this study, to reveal specific plant part features for 20 species from one family (Apiaceae), and to preserve the maximum information content, two approaches are suggested. In both cases, minimal raw data pretreatment, including rescaling of time and m/z axes and cutting off some uninformative regions, was applied. For the support vector machine (SVM) method, tensor unfolding was required, while neural networks (NNs) were able to work directly with squared heatmaps as input data. Moreover, five data augmentation variants are proposed, to overcome the typical problem of a lack of data. As a result, a comparable F1-score close to 0.75 was achieved by SVM and two employed NN architectures. Eight marker compounds belonging to chlorophylls, lipids, and coumarin apio-glucosides were tentatively identified as characteristic of their corresponding sample groups: roots, stems, leaves, and fruits. The proposed approaches are simple, information-saving and can be applied to a broad type of tasks in metabolomics.
INTRODUCTION Limited availability of individual standards is a bottleneck for quality control of functional foods and natural medicines. The use of standard mixtures or secondary standards is a possible alternative in this case. Earlier, an approach known as standardised reference extract (RE) strategy was introduced for HPLC-UV analysis of different plant materials; however, its application in HPLC-MS analysis has not been investigated. OBJECTIVE To establish an HPLC-MS-based RE method for determination of ginsenoside content in ginseng infusions using commercially available extract reference material of Panax quinquefolius L. RESULTS The developed HPLC-MS method was validated as precise (1.1%-9.4% intra-day variation; 1.6%-12.8% inter-day variation) and highly sensitive [limit of detection (LOD): 1-40 ng/mL; limit of quantification (LOQ): 4-120 ng/mL]. The stability of samples was satisfactory (5.7%-16.3%). The RE quantification method was compared with the external standard method, and the obtained difference was not significant, mostly in the range of 5%-10%. Matrix effects for the diluted samples of RE and ginseng infusions, determined via the standard addition method, were in the range of 85%-115% and 80%-126%, respectively, and were also positively correlated with the ginsenoside concentration. Eleven batches of ginseng infusions from different manufacturers were analysed using the established method. CONCLUSION The method for HPLC-MS-based ginsenoside quantification using RE as a secondary standard was established for the first time. The results of this study demonstrate that the application of the standardised RE strategy in HPLC-MS can minimise the matrix effect-related error in addition to the cost-effective quality control of herbal products, foods, and traditional medicines.
A combination of theoretical and experimental approaches was applied to determine the chromatographic rules of isomeric compounds’ behavior for preliminary identification. In gas chromatography-mass spectrometry (GC-MS), identification is performed by spectra matching, however, difficulties arise with isomeric compounds, which cannot be distinguished from each other without additional information. The thermodynamic characteristics of the adsorption of symmetric and asymmetric isomers of chlorophenylphenols, dimethoxybiphenyls, tri- and tetrachlorobiphenyls were determined using molecular statistical calculations. By-products in the chlorination of 4-hydroxybiphenyl were identified: 4-hydroxy-2,3′- and 3,2′-dichlorobiphenyls, 4-hydroxy-3,5,2′- and 2,3,6-trichlorobiphenyls. A developed theoretical approach was applied to predict the retention order of tri- and tetra-chlorobiphenyls. The GC-MS data and molecular statistical calculations made it possible to determine the main products of methoxybenzene dimerization as well as identify impurities. Thermodynamic parameters were received to describe the unusual retention behavior of epimers in reversed-phase high-performance liquid chromatography. Molecular descriptors were calculated to determine correlation with retention of both structural isomers and epimers. Descriptor combining surface area and partial charge information turned out to be useful in evaluating retention order for isomers.
New stationary phases for hydrophilic chromatography (HILIC) with functional layers formed by the multicomponent Ugi reaction have been obtained. Acetone, glycolic acid, ethyl isocyanacetate, and 3‑aminopropyl silica, which also act as adsorbent matrices, were used as reaction components. The study of the effect of solvent on the reaction yield showed that a high degree of coverage of the matrix was achieved when the reaction was carried out in ethanol. The new adsorbents’ chromatographic properties compared with the initial matrix were assessed using the Tanaka tests for hydrophilic phases and by studying the retention of polar analytes from various classes. The synthesized adsorbents have demonstrated high efficiency and selectivity in separating model mixtures of sugars, amino acids, and water-soluble vitamins in the HILIC mode.
In this work, simple, rapid and highly sensitive method of hazardous chemical 1,1-dimethylhydrazine (unsymmetrical dimethylhydrazine, UDMH) determination based on pre-column derivatization with unsubstituted aromatic aldehydes and reversed-phase high performance liquid chromatography-ultraviolet-tandem mass spectrometry (RP HPLC-UV-MS/MS) has been developed. Along with benzaldehyde, commercially available aromatic aldehydes, namely: 2-naphthaldehyde, 2-pyridinecarboxaldehyde, and 2-quinolinecarboxaldehyde, were used as derivatizing reagents in the analysis of hydrazines for the first time. The reactions were studied in a wide pH range by varying reaction time and other conditions. A slightly alkaline pH 9 was shown to be optimal for the derivatization of UDMH by aromatic aldehydes. The quantitative yield of derivatization products under the established conditions was confirmed by HPLC analysis with ampemmetric detection. For all studied reagents, wide linear ranges of concentrations (0.01-1000 mu g/L) in natural water samples were observed. The limits of detection for UDMH in natural water were in the 3.7-130 ng/L range. 2-Quinolinecarboxaldehyde was selected as the most appropriate reagent for HPLC-UV-MS/MS determination of UDMH. In case of using this reagent, the accuracy was in the range of 97-102%, and precision, expressed as RSD was less than 8%. The developed approach does not require laborious stages of pre-concentration and isolation of UDMH from natural water components.
The combination of Liquid Chromatography and Mass Spectrometry (LC-MS) is commonly used to determine and characterize biologically active compounds because of its high resolution and sensitivity. In this work we explore the interpretation of LC-MS data using multivariate statistical analysis algorithms to extract useful chemical information and identify clusters of similar samples. Samples of leaves from 19 plants belonging to the Apiaceae family were analyzed in unified LC conditions by high- and low-resolution mass spectrometry in a wide range scan mode. LC-MS data preprocessing was performed followed by statistical analysis using tensor decomposition in the form of Parallel Factor Analysis (PARAFAC); matrix factorization following tensor unfolding with principal component analysis (PCA), independent component analysis (ICA), non-negative matrix factorization (NMF); or unsupervised feature selection (UFS). The optimal number of components for each of these methods were found and results were compared using four different metrics: silhouette score, Davies-Bouldin index, computational time, number of noisy components. It was found that PCA, ICA and UFS give the best results across the majority of the criteria for both low- and high-resolution data. An algorithm for biomarker signal selection is suggested and 23 potential chemotaxonomic markers were tentatively identified using MS2 data. Dendrograms constructed by the methods were compared to the molecular phylogenic tree by calculating pixel-wise mean square error (MSE). Therefore, the suggested approach can support chemotaxonomic studies and yield valuable chemical information for biomarker discovery.
Retention time prediction in high-performance liquid chromatography (HPLC) is the subject of many studies since it can improve the identification of unknown molecules in untargeted profiling using HPLC coupled with high-resolution mass spectrometry. Lots of approaches were developed for retention time prediction in liquid chromatography for a different number of molecules considering various molecular properties and machine learning algorithms. The recently built large retention time data set of standard compounds from the Metabolite and Chemical Entity Database (METLIN) allows researchers to create a model that can be used for retention time prediction of small molecules with wide varieties of structures and physicochemical properties. The ability to predict retention times using the largest data set was studied for different architectures of deep learning models that were trained on molecular fingerprints, and SMILES (string representation of a molecule) represented as one-hot matrices. The best result was achieved with a one-dimensional convolutional neural network (1D CNN) that uses SMILES as an input. The proposed model reached the mean absolute error and the median absolute error equal to 34.7 and 18.7 s, respectively, which outperformed the results previously obtained for this data set. The pre-trained 1D CNN on the METLIN SMRT data set was transferred on five other data sets to evaluate the generalization ability.