For selected data set published by Russom et al. (Environ. Toxicol. Chem. 16, 948-967 (1997)) containing 704 organic molecules with measured acute aquatic toxicity data (96-h LC50 tests) we calculated data set of more than 1400 molecular descriptors by the Dragon 5.0 program. After we excluded descriptors that have almost constant values, and those having very low correlation with the logarithm of LC50 values on the training set, 623 descriptors remained and were used in the modeling process. Data set of molecules was randomly partitioned into the training and test set containing 560 and 144 molecules, respectively. We developed and compared two kinds of ensemble of both linear and nonlinear multi-regression models (1) normal ensembles and (2) ensembles obtained by the clustering of molecules according to their similarity (clustered ensembles). Clustering of molecules was performed by calculating their Euclidian distances in normalized descriptor space. In this method, the final model was developed only on those molecules from the training set that are close (measured using Euclidian distance in normalized descriptor space) to the selected molecule from the test set. Although results obtained by normal ensembles are very good (e.g. nonlinear ensemble of 8-descriptor models: r(2) = 0.82, s = 0.54 (training set), s(test) = 0.80), significant improvement is obtained by taking into account clustering of molecules in development of ensembles of linear models (e.g. 200 10-descriptor models in ensemble: r(2) = 0.87, s(train) = 0.45 (training set), s(test) = 0.76; or for 200 simpler models having 7-descriptor models in ensemble r(2) = 0.83, s(train) = 0.53 (training set), s(test) = 0.77). These results clearly indicate that the use of information about similarity between molecules can improve structure-toxicity models, and we also expect that this could be valid generally.
Several quantitative structure-activity studies for this data set containing 107 HEPT derivatives have been performed since 1997, using the same set of molecules by (more or less) different classes of molecular descriptors. Multivariate Regression (MR) and Artificial Neural Network (ANN) models were developed and in each study the authors concluded that ANN models are superior to MR ones. We re-calculated multivariate regression models for this set of molecules using the same set of descriptors, and compared our results with the previous ones. Two main reasons for overestimation of the quality of the ANN models in previous studies comparing with MR models are: (1) wrong calculation of leave-one-out (LOO) cross-validated (CV) correlation coefficient for MR models in Luco et al., J. Chem. Inf. Comput. Sci. 37 392-401(1997), and (2) incorrect estimation/interpretation of leave-one-out (LOO) cross-validated and predictive performance and power of ANN models. More precise and fairer comparison of fit and LOO CV statistical parameters shows that MR models are more stable. In addition, MR models are much simpler than ANN ones. For real testing the predictive performance of both classes of models we need more HEPT derivatives, because all ANN models that presented results for external set of molecules used experimental values in optimization of modeling procedure and model parameters.
Flavonoid derivatives are very important class of bioactive compounds often used in drug design. Several quantitative structure-activity studies for 104 flavonoid derivatives and their inhibition of p56lck Protein Tyrosine Kinase (PTK) were performed using different classes of molecular descriptors. However, published models are not of high accuracies - the best mode achieves r = 0.81 (correlation coefficient) and s = 0.43 (standard error of estimate) for the complete data set (104 flavonoids). In our recent study old results were outperformed by using novel sets of molecular descriptors computed by (1) the DRAGON 4.0 program and (2) the CODESSA 2.21 program. Somewhat better model we obtained using four DRAGON descriptors r = 0.85, s = 0.38, and s(cv) = 0.39 (leave-one-out cross-validated standard error of estimate). In this study we present further improvement of models by using molecular descriptors that are based on autocorrelation functions weighted by, different atomic properties that were computed by the ADRIANA.code program. Data set was encoded as SMILES and converted to 3D structures (SD files) by the CORINA program (www2.chemie.uni-erlangen.de/software/corina/), and more than 700 descriptors were calculated by the ADRIANA.code program. The selection of the most significant molecular descriptors into Multivariate Linear Regression (MLR) models containing 1-3 descriptors were performed by the CROMRsel program. The best model containing three descriptors had s = 0.29 and s(cv) = 0.31 for 104 molecules. The best model contains 2D autocorrelation descriptor of order 11 weighted by the sigma atom charge, and two 3D autocorrelation functions of orders 9 and 4 weighted by the total charge, and sigma electronegativity, respectively. This is promising result showing that in data sets of molecules that have large common portion of their structures, descriptors based on autocorrelation are most significant ones.
In this study we want to test whether a simple modeling procedure used in the field of QSAR/QSPR can produce simple models that will be, at the same time, as accurate as robust Neural Network Ensemble (NNE) ones. We present results of application of two procedures for generating/selecting simple linear and nonlinear multiregression (MR) models: (1) method for selecting the best possible MR models (named as CROMRsel) and (2) Genetic Function Approximation (GFA) method from the Cerius2 program package. The obtained MR models are strictly compared with several NNE models. For the comparison we selected four QSAR data sets previously studied by NNE (Tetko et al. J. Chem. Inf. Comput. Sci. 1996, 36, 794-803. Kovalishyn et al. J. Chem. Inf. Comput. Sci. 1998, 38, 651-659.): (1) 51 benzodiazepine derivatives, (2) 37 carboquinone derivatives, (3) 74 pyrimidines, and (4) 31 antimycin analogues. These data sets were parameterized with 7, 6, 27, and 53 descriptors, respectively. Modeled properties were anti-pentylenetetrazole activity, antileukemic activity, inhibition constants to dihydrofolate reductase from MB1428 E. coli, and antifilarial activity, respectively. Nonlinearities were introduced into the MR models through 2-fold and/or 3-fold cross-products of initial (linear) descriptors. Then, using the CROMRsel and GFA programs (J. Chem. Inf. Comput. Sci. 1999, 39, 121-132) the sets of I (I < or = 8, in this paper) the best descriptors (according to the fit and leave-one-out correlation coefficients) were selected for multiregression models. Two classes of models were obtained: (1) linear or nonlinear MR models which were generated starting from the complete set of descriptors, and (2) nonlinear MR models which were generated starting from the same set of descriptors that was used in the NNE modeling. In addition, the descriptor selection method from CROMRsel was compared with the GFA method included in the QSAR module of the Cerius2 program. For each data set it has been found that the MR models have better cross-validated statistical parameters than the corresponding NNE models and that CROMRsel selects somewhat better MR models than the GFA method. MR models are also much simpler than NNEs, which is the important surprising fact, and, additionally, express calculated dependencies in a functional form. Moreover, MR models were shown to be better than all other models obtained by different methods on the same data sets ("old" multivariate regressions, functional-link-net models, back-propagation neural networks, genetic algorithm, and partial least squares models). This study also indicated that the robust NNE models cannot generate good models when applied on small data sets, suggesting that it is perhaps better to apply robust methods (like NNE ones) on larger data sets.