Dissolved organic matter (DOM) acts as an inducer as well as an inhibitor in the photodegradation of 4-(N,N-dimethylamino)benzonitrile (DMABN), which undergoes both photodegradation by the excited triplet states of chromophoric DOM ((CDOM)-C-3*), and inhibition upon back reduction by the antioxidant moieties of DOM itself. On average, water with dissolved organic carbon (DOC) > 2.3 mgC L-1 would inhibit DMABN photodegradation by over 50 %. From the available data of [(CDOM)-C-3*] steady-state concentrations and the levels of DOC in lake water on a quasi-global scale (60 degrees S-60 degrees N latitude range), we were able to numerically assess the photochemical lifetimes of DMABN, here reported as year averages. Lifetimes of DMABN in lake water would amount to a couple of months or longer, and they would be the shortest in the tropical belt. Results of numerical calculations were interpolated using a suitable model function, with which an analytical equation was derived that relates DMABN lifetimes with lake depth, DOC, and latitude. For the first time to our knowledge, a general equation is here proposed to assess the photodegradation kinetics of a pollutant on a (quasi-)global scale.
Nowadays, the large number of measurable variables has considerably increased the complexity of data. In the framework of the decision-making process, this leads to the need of adequate tools to set priorities and rank the available options. Ordering is one of the possible ways to analyse multivariate data, which provides an overview of the relationships among the elements of a system. The Multi-Criteria Decision Making (MCDM) encompasses a broad set of methods designed to set priority-based lists of alternatives based on multiple criteria, which support decision problems. Among the most widely adopted techniques, TOPSIS, dominance-based approaches, the Analytic Hierarchy Process (AHP), and Copeland scores represent some of the classical methodologies in both theoretical research and applied decision analysis. Among the dominance-based approaches, an effective MCDM method is the Power-Weakness Ratio (PWR), which generates a tournament table (i.e., the pairwise comparison matrix) from a data matrix with a varying number of samples (i.e., alternatives to be compared) and variables (i.e., the criteria for pairwise comparisons), weighted according to their relative importance in determining the final ranking. In this study, a variant of the classical Power-Weakness Ratio is presented, significantly modifying the way the tournament table is obtained. The method, called smoothed Power-Weakness Ratio (sPWR), takes into account the dominance degree of the alternatives in each pairwise comparison exploiting the differences between the criterion values. The rationale behind the method is described by the aid of an illustrative example on a simple benchmark dataset with known reference ranking of the samples. The main advantage of the new method over PWR is that its tournament table is much more informative and sensitive to the original data values than the classical pairwise comparison matrix. A multivariate comparison with other classical MCDM methods, performed on several diverse datasets, demonstrated that the results obtained by sPWR were quite similar to those obtained by Copeland Score and TOPSIS with range scaling. However, sPWR showed a higher tendency toward generating full rankings with an enhanced ability to remove ties in the pairwise comparisons.
Photochemical modeling was used to evaluate the photodegradation kinetics of sulfamethoxazole (SMX) in sunlit natural freshwater environments, utilizing available data on SMX absorption spectrum, direct photolysis quantum yield, and second-order reaction rate constants with photochemically produced reactive intermediates (PPRIs). The anionic form of SMX is considered here, which prevails in the majority of surface-water conditions (pH > 5.6). SMX would be photodegraded quite rapidly in well-mixed waters under sunny weather, with direct photolysis serving as the predominant pathway in most cases. Among the water parameters, depth and the dissolved organic carbon (DOC) primarily influence the photodegradation of SMX, which would be faster in shallow waters with low DOC. This feature enabled the extension of SMX photochemical lifetime prediction to lakes located within the 60°S-60°N latitude belt, where data on DOC, water depth and sunlight irradiance are available. The fastest photodegradation of SMX (annual average) is expected in the tropical belt, where the combination of high irradiance and low DOC would be most favorable for direct photolysis. With fair weather and good water mixing, typical SMX lifetimes would range from days to some weeks depending on the specific conditions. The lifetimes predicted in shallow waters are consistent with several literature reports of SMX photodegradation kinetics in the laboratory.
Liquid chromatography (LC) coupled to mass spectrometry (MS) is a powerful and versatile technique with several applications in analytical chemistry such as the untargeted analysis of complex mixtures of plant food bioactive compounds. In this framework, identification of compounds mainly relies on mass and fragmentation patterns derived from MS spectra of available databases. However, to develop automated tools for a rapid compound identification, prior knowledge of retention times (RTs) has been demonstrated to be a valid support to MS data to reduce the number of possible candidate structures. Unlike experimental methods, which are time-consuming and limited to a few classes of compounds, data-driven computational approaches can predict retention times for a wide range of compounds using only their molecular structures and physicochemical properties. In this research, Genetic Algorithms (GAs) coupled to Multiple Linear Regression (MLR) were applied to select the most relevant molecular descriptors to establish quantitative structure-retention relationships (QSRRs) aimed at predicting the retention times of plant food bioactive compounds across three different LC chromatographic systems. The statistical parameters showed model robustness and satisfactory predictive ability. Particular attention was paid to measuring the uncertainty of predictions and assessing their reliability based on the model applicability domain. Interpretation of the selected molecular descriptors provided valuable insight into the separation mechanism. Finally, the developed models were applied to predict the unknown retention times, for the three studied LC chromatographic systems, of a large library of plant food bioactive compounds, which were made freely available to further assist the research in the field of natural products.
Natural products are a diverse class of compounds with promising biological properties, such as high potency and excellent selectivity. However, they have different structural motifs than typical drug-like compounds, e.g., a wider range of molecular weight, multiple stereocenters and higher fraction of sp3-hybridized carbons. This makes the encoding of natural products via molecular fingerprints difficult, thus restricting their use in cheminformatics studies. To tackle this issue, we explored over 30 years of research to systematically evaluate which molecular fingerprint provides the best performance on the natural product chemical space. We considered 20 molecular fingerprints from four different sources, which we then benchmarked on over 100,000 unique natural products from the COCONUT (COlleCtion of Open Natural prodUcTs) and CMNPD (Comprehensive Marine Natural Products Database) databases. Our analysis focused on the correlation between different fingerprints and their classification performance on 12 bioactivity prediction datasets. Our results show that different encodings can provide fundamentally different views of the natural product chemical space, leading to substantial differences in pairwise similarity and performance. While Extended Connectivity Fingerprints are the de-facto option to encoding drug-like compounds, other fingerprints resulted to match or outperform them for bioactivity prediction of natural products. These results highlight the need to evaluate multiple fingerprinting algorithms for optimal performance and suggest new areas of research. Finally, we provide an open-source Python package for computing all molecular fingerprints considered in the study, as well as data and scripts necessary to reproduce the results, at https://github.com/dahvida/NP_Fingerprints .
Clustering is an unsupervised machine learning methodology widely used in several sciences to find groups of similar patterns in complex data. The results generated by clustering algorithms generally depend on user-defined input parameters such as the number of expected clusters, which can have a great impact on the homogeneity of the identified clusters.Clustering validity indices (CVIs) are an effective method for determining the optimal number of clusters that best fit the natural partition of a dataset. They do not require any underlying assumption nor a priori knowledge about the true dataset structure. Since 1965, many cluster validity indices have been proposed in the literature and used in several different applications.In this paper, the performance of 68 cluster validity indices was evaluated on 21 real-life research and simulated datasets. CVIs were compared on the same partition for each dataset, which was searched for by the k-means clustering algorithm. Multivariate chemometric methods were applied to disclose mutual relationships among the indices and to select those that are more effective in terms of accuracy and reliability.
Approaches of high-level data fusion, also known as consensus, combine predictions of individual models to increase reliability and overcome limitations of single models. Consensus strategies are frequently applied in the framework of Quantitative Structure - Activity Relationships (QSARs) to reduce the uncertainties in the prediction of molecular activities and provide better accuracy of the model outcomes. However, specific regions of the chemical space may systematically be associated with low accuracy and even consensus modelling cannot improve prediction reliability through the multiple outcomes of individual models.In this study, a new heuristic metric to assess the degree of accuracy of consensus predictions in the chemical space is proposed. This metric can assist the mapping of reliability in prediction and enhance the delineation of a safe zone, where consensus predictions are expected to have better accuracy. The new metric is calculated by kernel-based potential functions and it can be used in the framework of both classification and regression consensus modelling. Four case studies, including extensive datasets for consensus modelling, were used to test the proposed approach.Results demonstrated that a potential can be associated with regions of the chemical space as a function of accuracy of consensus modelling and it can be used to enable the mapping of reliability in prediction and the definition of specific regions where predictions are expected to be more reliable.
The capacity to discriminate safe from dangerous compounds has played an important role in the evolution of species, including human beings. Highly evolved senses such as taste receptors allow humans to navigate and survive in the environment through information that arrives to the brain through electrical pulses. Specifically, taste receptors provide multiple bits of information about the substances that are introduced orally. These substances could be pleasant or not according to the taste responses that they trigger. Tastes have been classified into basic (sweet, bitter, umami, sour and salty) or non-basic (astringent, chilling, cooling, heating, pungent), while some compounds are considered as multitastes, taste modifiers or tasteless. Classification-based machine learning approaches are useful tools to develop predictive mathematical relationships in such a way as to predict the taste class of new molecules based on their chemical structure. This work reviews the history of multicriteria quantitative structure-taste relationship modelling, starting from the first ligand-based (LB) classifier proposed in 1980 by Lemont B. Kier and concluding with the most recent studies published in 2022.
Multitask learning allows to model multiple tasks simultaneously through information sharing. In the context of quantitative structure–activity relationships and computational toxicology, multitask learning is gaining more and more interest, owed to its potential to improve the predictive performance of underrepresented tasks and to predict the multi-property profile of molecules. In this chapter, after introducing the multitask problem formulation, we present a hands-on tutorial on multitask neural networks.
With the continuous growth of the real and virtual chemical space, efficient computer-assisted methods are required to discover new substances with desired properties and/or predict properties of interest for untested molecules. These methods rely on the principle that the physicochemical and biological properties of compounds are the effects of their structural characteristics. Therefore, the starting point of any chemo- and bioinformatics application is the conversion of a symbolic representation of the molecular structure into numerical information through the calculation of molecular descriptors. Molecular descriptors encode a wide variety of specific molecular features with a different effect on experimental properties and impact on the perceived chemical similarity between molecules. The choice of molecular descriptors is crucial in determining the chemical space representation and computational modeling outcomes. After introducing the fundamental concepts of molecular descriptors in the current epistemological framework, this chapter reviews some of the well-known classical molecular descriptors and fingerprints.
The interest in multitask and deep learning strategies has been increasing in the last few years, in application to large and complex dataset for quantitative structure-activity relationship (QSAR) analysis. Multitask approaches allow the simultaneous prediction of molecular properties that are related, through information sharing, whereas deep learning strategies increase the potential of capturing nonlinear relationships. In this work, we compare the binary classification capability of multitask deep and shallow neural networks to single-task strategies used as benchmark (i.e., as k-nearest neighbours, N-nearest neighbours, random forest and Naive Bayes), as well as multitask supervised self-organizing maps. Comparison was carried out with an extended QSAR dataset containing annotations of molecular binding, agonism and antagonism activity on 11 nuclear receptors, for a total of 14,963 molecules, divided into training and test sets and labelled for their bioactivity on at least one of 30 binary tasks. Additional 304 chemicals were used as external evaluation set to further validate models. Although no approach systematically overperformed the others, task-specific differences were found, suggesting the benefit of multitask learning for tasks that are less represented. On average, some of the single-task approaches and multitask deep learning strategies had similar performances. However, the latter can have advantages, such as a simpler management of predictions and applicability domain assessment for future samples. On the other hand, the parameter tuning required by neural networks are generally time expensive suggesting that the modelling strategy should be evaluated case by case.
The study concerns the photodegradation of the antidepressant escitalopram (ESC), the S-enantiomer of the citalopram raceme, both in ultrapure and surface water, considering the contribution of indirect photolysis through the presence of nitrate and bicarbonate. The effect of nitrate and bicarbonate concentrations was investigated by full factorial design, and only the nitrate concentration resulted in having a significant effect on the degradation. The kinetics of ESC photodegradation is the pseudo-first-order (half-life = 62.4 h in ultrapure water and 48.4 h in lake water). The generation of transformation products (TPs) was monitored through a developed and validated HPLC-MS/MS method. Fourteen TPs were identified in ultrapure water (one of them, at m/z 261, for the first time) and other two TPs at m/z 327 (found for the first time in this study) were identified only in presence of a nitrate. Several TPs were the same as those formed during the photodegradation of citalopram. The photodegradation pathway of ESC and its mechanism of degradation in water is proposed. The method was applied successfully to the analyses of surface water samples, in which a few dozen of ng L−1 of ESC was determined together with the presence of TP2, TP5 and TP12. Finally, a preliminary in silico evaluation of the toxicological profile and environmental behavior of TPs by computational models was carried out; two TPs (TP4 and TP10) were identified as of potential concern, as they were predicted mutagenic by Ames test model.
Mass spectrometry (MS) is widely used for the identification of chemical compounds by matching the experimentally acquired mass spectrum against a database of reference spectra. However, this approach suffers from a limited coverage of the existing databases causing a failure in the identification of a compound not present in the database. Among the computational approaches for mining metabolite structures based on MS data, one option is to predict molecular fingerprints from the mass spectra by means of chemometric strategies and then use them to screen compound libraries. This can be carried out by calibrating multi-task artificial neural networks from large datasets of mass spectra, used as inputs, and molecular fingerprints as outputs. In this study, we prepared a large LC-MS/MS dataset from an on-line open repository. These data were used to train and evaluate deep-learning-based approaches to predict molecular fingerprints and retrieve the structure of unknown compounds from their LC-MS/MS spectra. Effects of data sparseness and the impact of different strategies of data curing and dimensionality reduction on the output accuracy have been evaluated. Moreover, extensive diagnostics have been carried out to evaluate modelling advantages and drawbacks as a function of the explored chemical space.
Minimum Spanning Tree (MST) is a well-known clustering algorithm that provides a graphical tree representation of the objects in a data set by exploiting local information to link each pair of similar objects. The a-posteriori analysis of this tree in terms of nodes and edges provides the basis to derive simple classifiers, namely semi-supervised classification approaches based on the minimum spanning tree approach. In this work, we propose different metrics to evaluate the MST ability to group objects of the same a-priori known classes. The classification capability of the proposed approach, using 13 different distance measures, was compared with that of classical supervised classification approaches such as N-Nearest Neighbour (N3), Binned Nearest Neighbour (BNN), Partial Least Squares-Discriminant Analysis (PLS-DA), K-Nearest Neighbour (KNN), exponentially weighted K-Nearest Neighbour (wKNN) and Support Vector Machine with radial functions (SVM-RBF) on 31 data sets. The proposed approach resulted to be competitive and comparable with the considered classical supervised classification methods. Finally, we analysed the role of the 13 different measures in terms of performance and percentage of not-assigned objects.
Nuclear receptors (NRs) are involved in fundamental human health processes and are a relevant target for toxicological risk assessment. To help prioritize chemicals that can mimic natural hormones and be endocrine disruptors, computational models can be a useful tool.1,2 In this work we i) created an exhaustive collection of NR modulators and ii) applied machine learning methods to fill the data-gap and prioritize NRs modulators by building predictive models.
BACKGROUND:Humans are exposed to tens of thousands of chemical substances that need to be assessed for their potential toxicity. Acute systemic toxicity testing serves as the basis for regulatory hazard classification, labeling, and risk management. However, it is cost- and time-prohibitive to evaluate all new and existing chemicals using traditional rodent acute toxicity tests. In silico models built using existing data facilitate rapid acute toxicity predictions without using animals. OBJECTIVES:The U.S. Interagency Coordinating Committee on the Validation of Alternative Methods (ICCVAM) Acute Toxicity Workgroup organized an international collaboration to develop in silico models for predicting acute oral toxicity based on five different end points: Lethal Dose 50 (LD50 value, U.S. Environmental Protection Agency hazard (four) categories, Globally Harmonized System for Classification and Labeling hazard (five) categories, very toxic chemicals [LD50 (LD50≤50mg/kg)], and nontoxic chemicals (LD50>2,000mg/kg). METHODS:An acute oral toxicity data inventory for 11,992 chemicals was compiled, split into training and evaluation sets, and made available to 35 participating international research groups that submitted a total of 139 predictive models. Predictions that fell within the applicability domains of the submitted models were evaluated using external validation sets. These were then combined into consensus models to leverage strengths of individual approaches. RESULTS:The resulting consensus predictions, which leverage the collective strengths of each individual model, form the Collaborative Acute Toxicity Modeling Suite (CATMoS). CATMoS demonstrated high performance in terms of accuracy and robustness when compared with in vivo results. DISCUSSION:CATMoS is being evaluated by regulatory agencies for its utility and applicability as a potential replacement for in vivo rat acute oral toxicity studies. CATMoS predictions for more than 800,000 chemicals have been made available via the National Toxicology Program's Integrated Chemical Environment tools and data sets (ice.ntp.niehs.nih.gov). The models are also implemented in a free, standalone, open-source tool, OPERA, which allows predictions of new and untested chemicals to be made. https://doi.org/10.1289/EHP8495.
Multivariate regression is a fundamental supervised chemometric approach that defines the relationship between a set of independent variables and a quantitative response. It enables the subsequent prediction of the response for future samples, thus avoiding its experimental measurement. Regression approaches have been widely applied for data analysis in different scientific fields. In this paper, we describe the regression toolbox for MATLAB, which is a collection of modules for calculating some well-known regression methods: Ordinary Least Squares (OLS), Partial Least Squares (PLS), Principal Component Regression (PCR), Ridge and local regression based on sample similarities, such as Binned Nearest Neighbours (BNN) and k-Nearest Neighbours (kNN) regression methods. Moreover, the toolbox includes modules to couple regression approaches with supervised variable selection based on All Subset models, Forward Selection, Genetic Algorithms and Reshaped Sequential Replacement. The toolbox is freely available at the Milano Chemometrics and QSAR Research Group website and provides a graphical user interface (GUI), which allows the calculation in a user-friendly graphical environment.