
A new 3D-QSAR method based on the novel molecular dynamics methodology, Active Site Pressurization (ASP), has been validated using two cyclin-dependent kinase 2 data sets containing 65 purines and 91 oxindoles. ASP allows the construction of cavity casts that represent the maximal energetically feasible 3D distortion of protein binding sites potentially achievable by induced fit upon binding of ligands. The ASP-QSAR method entails many components of traditional 3D-QSAR strategies but additionally correlates the biological activity of ligand sets with features of ASP-derived binding site cavity casts, thus taking target protein flexibility into account implicitly. Both of the data sets used to validate the ASP-QSAR method resulted in QSAR models that were of exceptional quality and predictivity. A non-cross-validated variance coefficient (R-2) between 0.959 and 0.99 and a cross-validated variance coefficient (Q(2)) of between 0.927 and 0.929 were obtained for these ASP-QSAR models.
Structure-based Quantitative Structure-Activity Relationship (QSAR) studies were performed on the Human Immunodeficiency Virus Reverse Transcriptase (HIV-RT) inhibitors with Diaryltriazines (DATAs) and Diarylpyrimidines (DAPYs) using Comparative Molecular Field Analysis (CoMFA) and Comparative Molecular Similarity Indices Analysis (CoMSIA) implemented in the SYBYL software packages. From the Xray structure of dapivirine, 38 training set molecules were sketched and minimized using the MMFF94 force field. In structure-based QSAR, all the HIV-RT complexes were subjected to subset minimization and aligned against a fixed point of the enzyme. The best CoMFA and CoMSIA results presented cross-validated values (q(2)) of 0.640 and 0.663, and non-cross-validated values (r(2)) of 0.976 and 0.932, respectively. Contour map analysis enables the identification of crucial interactions between the enzyme and inhibitors, which can be used further to design new HIV-RT inhibitors.
In this article the calculation scheme of new molecular descriptors on the basis of symbiosis of the informational field model and simplex representation of molecular structure (so-called, simplex-informational descriptors) is introduced. The following advantages of the proposed descriptors are demonstrated: i) high-sensitivity to molecular structure changes; ii) good mechanistic interpretation; and iii) the ability to produce well-fitted, robust and predictive QSAR models. The efficiency of the method is also demonstrated on the example of QSAR analysis of angiotensin converting enzyme (ACE) inhibitors, acetylcholinesterase (AChE) inhibitors, and ligands of 5-HTIA receptors. QSAR tasks have been solved using the PLS-method. Resulting 2D models are obtained using simplex-informational and original simplex descriptors and were also compared with descriptors generated by other QSAR approaches (COMFA, COMSIA, EVA, HQSAR). The advantage of the developed method over others are shown by the comparison of different statistical characteristics.
As part of the ATP-binding cassette transporter superfamily P-glycoprotein (ABCB1) acts as xenotoxic exporter and consequently is strongly involved in multidrug resistance (MDR) and drug-drug interactions. In this work we focus on our in-house developed SIBAR approach for prediction of ABCB1 substrates. SIBAR values were calculated on basis of three different descriptor sets: 2D-MOE descriptors, VSA descriptors and 3D Autocorrelation vectors using in total four reference sets. In order to compare linear with non-linear classification methods we used binary QSAR and a support vector machine (SVM), respectively. Results demonstrate that with 2D-MOE and VSA-descriptors prediction of non substrates performs better, whereas autocorrelation vectors show higher accuracy for substrates. With respect to the different reference sets used in this case selection on basis of maximum diversity yielded better results than a set derived from the training set compounds. In general, the models show distinct differences in their performance depending on the combination of method and descriptor type.
The cytotoxicity of 19 N-aryl-substituted hydroxamic acids has been tested in vitro towards human breast cancer MCF-7 cell lines by MTT assay. The IC(50) values were found to be in the range from 61.94 to 337.54 mu M. A total of 18 out of 19 molecules had higher inhibitory activities than hydroxyurea against MCF-7 cell. Five compounds with IC(50) values in micromolar range were 3- to 5-folds more potent than hydroxyurea (IC(50)= 307.15 mu M). By partial least squares (PLS) regression, 2D-QSAR model reported herein provide interesting insight in understanding hydrophobic, electronic, and structural requirements of antitumor activity among these set of the compounds. The cross-validated Q(cum)(2) values for optimal PLS model of hydroxamic acids is above 0.638 (remarkably higher 0.50), indicating good predictive abilities for log 1/IC(50) values of hydroxamic acids (HAS). The k-nearest neighbor molecular field analysis (kNN-MFA) approach was used to generate three-dimensional quantitative structure-activity relationship (3D-QSAR) models for these sets of molecules. Statistically stepwise variable selection k-nearest neighbor molecular field analysis (SW-kNN-MFA) model is comparatively better as compared to the other two (i.e. simulated annealing k-nearest neighbor molecular field analysis, SA-kNN-MFA, and genetic algorithm k-nearest neighbor molecular field analysis, GA-kNN-MFA) with respect to both the internal (q(2)=0.7461) as well as external (pred_r(2)=0.6107) model validation and correctly predicts activity of ca. 74.61% and ca. 61.07% for the training and test set, respectively. It uses one steric and one electrostatic fields along with its 3k nearest neighbor (k=3) to evaluate the activity of new molecules. The developed SW-kNN-MFA model field plot indicated that the positive steric and electric potential are favorable for the increase in the activity and hence more bulky substituent at 5-position of phenyl ring connected at amide group and less electronegative substituent at 3-position of phenyl ring connected to carboxyl group are favourable for the increase in the potency of the molecules.
The gas chromatography retention indices of 168 pesticides were used to construct a robust quantitative structure-retention relationship (QSRR) model. After outlier detection by Cook's influence measurement, the remaining compounds were subjected to two different modeling strategies. The first one was stepwise multiple linear regression (stepwise-MLR). Results of this method revealed that 81.7 percent of variances of the response could be explained by the model. The other strategy was kernel orthogonal projection to latent structure (KOPLS). R-2 and RMSE values for the prediction set established by Monte Carlo cross validation of the KOPLS were 0.906 and 0.093, respectively. Y-randomization technique was used to assess the chance correlation in the developed models. This technique indicated that there was no chance correlation in the KOPLS constructed model. From the results of this work one may conclude that KOPLS Is superior over the stepwise-MLR in term of predictability and can be used as an alternative method in quantitative structure - activity/- retention relationship (QSAR/QSRR) studies.
During the years the National Cancer Institute (NCI) accumulated an enormous amount of information through the application of a complex protocol of drugs screening involving several tumor cell lines, grouped into panels according to the disease class. The Anti-cancer Agent Mechanism (ACAM) database is a set of 122 compounds with anti-cancer activity and a reasonably well known mechanism of action, for which are available drug screening data that measure their ability to inhibit growth of a panel of 60 human tumor lines, explicitly designed as a training set for neural network and multivariate analysis. The aim of this work is to adapt a methodology (previously developed for the analysis of DNA minor groove binders) for the analysis of NCI ACAM database, using Principal Component Analysis (PCA) and QSAR/QSPR for the prediction of the mechanism of action of anti-cancer drugs. The entire database was splitted in a training set of 60 structures and a test set of 48 ones, and each set was expressed in form of a matrix on which further procedures were performed. Three statistical parameters were calculated: First Attempt of Prediction (FAP) expresses the percentage of correct predictions at first attempt, Total Attempt of Prediction (TAP) expresses the total percentage of correct predictions across all the three attempts, Non-Classified (NC) expresses the percentage of compounds whose mechanism of action has failed to be predicted. The predictive ability of this approach is variable, but the results obtained are generally good; using 50% Growth Inhibiting concentration (GI50) values as training data, we were able to assign a correct mechanism of action with a good degree of reliability (more than 79%).
Prediction accuracy of in silico methods for physicochemical and ADMET properties of drugs is an actual matter of controversial discussions. With a particular concern on log P prediction methods, we discuss here, how understanding the limitations of methods, their applicability domains and their prediction accuracies, as well as the use of local models can help to establish accurate and meaningful in silico predictions.
The predictive performance of five different pK(a) prediction tools (ACDpKa, Epik, Marvin pKa, Pallas pKa, and VCCpKa) was investigated on the 248-membered Gold Standard dataset. We found VCC as the most predictive, high throughput pK(a) predictor. However since VCC calculates pK(a) for the most acidic or basic group only we concluded that ACD and Marvin are in fact the method of choice for medicinal chemistry applications. Analyzing the common outliers we identified guanidines, enolic hydroxyl groups and weak acidic NHs as most problematic moieties from prediction point of view. Our results obtained on the high quality, homogenous Gold Standard dataset could be useful for end-users selecting a suitable solution for pK(a) prediction.
This study describes the ligand based as well as structure based molecular modeling and virtual screening of selective tumor necrosis factor-a converting enzyme (TACE) inhibitors. In ligand based molecular modeling, two statistically reliable pharmacophore models HypoA1 and HypoB1 were generated using a same training set of 22 molecules. HypoA1 consists of two hydrogen bond acceptor and three hydrophobic groups whereas HypoB1 consists of one hydrogen bond donor, one ring aromatic and three hydrophobic groups. Virtual screening was performed with both models in in-house database of 1.2 million molecules. To remove non selective hits from screened molecules, a counter pharmacophore was generated using inhibitors of MMP-1, an important enzyme involved in musculoskeletal degradation. In structure based molecular modeling, docking analysis was performed to explore the important interactions between ligands and protein. On comparison, HypoA1 and HypoB1 were found to be complementing with results of docking analysis suggesting high reliability of both models for their use ill virtual screening/designing of new molecule.
In this work, the receptor-bound conformation of nine vasopressin (CYFQNCPRG-NH2, AVP) analogs substituted in positions 2 or 3 with 2-aminoindane-2-carboxylic acid have been investigated using molecular modeling methods. The synthesis and functional assays of the analogs have been recently described. For comparison, the molecular dynamics of the selected analogs has been conducted. We have observed the relationship between the value of valence angle between the aromatic rings in positions 2 and 3 and the biological activity of the analogs. Both, in restricted space of the receptor cavities and unlimited continuous water environment, the same valence angles prevail, thus seem to be energetically more favorable for the particular peptides. Moreover, the residues responsible for analogs binding to the receptors have been identified.
Steroid sulfatase (STS) is the steroidogenic enzyme responsible for the hydrolysis of different sulfated steroids into their corresponding hydroxylated forms. This enzyme attracts our attention for its potential role in the growth of hormone-dependent breast and prostate tumors by the transformation of inactive sulfated precursors (which are very abundant in the blood) into active sex steroids. In order to identify the parameters responsible for good affinity with the active enzyme site and thus producing a good reversible STS inhibitor, we have built a quantitative structure-activity relationship (QSAR) model by using MDL-QSAR software which analyzes the molecules through more than 400 molecular descriptors. A total of 65 derivatives in position 17 alpha of estradiol with their corresponding IC50 values were used to create our OSAR model. The linear regression converged through an optimization process to a relatively simple equation described by 4 molecular descriptors (Log P, nelem, kappa(0) and kappa alpha(3)). Virtual screening of approximately 200 molecules then enabled us to direct the synthesis of new reversible STS inhibitors.
This paper reviews the articles published in Volumes 4-27 of the journals Quantitative Structure-Activity Relationships and of QSAR & Combinatorial Science, focusing on the articles published in the journals, citations to those articles, the most productive authors and countries, and the relationship of the journals to the more general chemical literature.
Comparative Molecular Field Analysis (CoMFA) is a popular Three-Dimensional Quantitative Structure-Activity Relationship (3D-QSAR) method. The effect of varying molecular fields [CoMFA vs. Comparative Molecular Similarity Indices Analysis (CoMSIA)], lattice spacing and analysis options on predictive performance was assessed based on cross validated R(2) values of 30 datasets taken from the literature. The CoMFA method results in statistically significantly higher cross validated R(2) values compared to COMSIA when only steric and electrostatic fields are used. When the hydrophobic molecular field is included in the CoMSIA analysis (as is most commonly the case) the difference between CoMFA and CoMSIA is no longer statistically significant. Addition of hydrogen bond field does not improve CoMFA predictivity. Although there was a trend towards increased predictivity with decreased lattice spacing, this was not statistically significant. Altering the default filtering criteria and cut-off values for Lennard Jones and Coulomb fields also did not generally result in a statistically significant effect on predictive performance.
D-Alanyl - D-alanine ligase is an enzyme which catalyzes the dimerization of D-alanine, and, as such, has an essential role in bacterial cell wall biosynthesis. It has been shown that inhibition Of D-alanyl - D-alanine ligase prevents bacterial growth. D-Alanyl D-alanine ligase represents therefore a viable antimicrobial target. The 3D structure of this enzyme complexed with a phosphinophosphate inhibitor has been reported, which allows for structure-based design studies. Four softwares (LUDI, MCSS, Autodock, and Glide) developed either for fragment or full-molecule docking were compared and scored for their ability to position in the active site four prototypic ligands: two inhibitors, i.e. a phosphinophosphate derivative and D-cycloserine, D-alanine and D-alanyl - D-alanine. Best performances were obtained with Glide and MCSS. A short series of novel derivatives based on a 2-phenylbenzoxazole scaffold was designed de novo on the basis of computational data. The best compound was found to fully inhibit the D-alanyl D-alanine ligase of E. faecalis with an IC50 of 400 mu M.
Pirinixic acid is a moderate agonist of both the alpha and the gamma subtype of the peroxisome proliferator activated receptor (PPAR). Previously, we have shown that alpha-alkyl substitution leads to balanced low micromolar-active dual agonists of PPAR alpha and PPAR gamma. Taking alpha-hexyl pirinixic acid as a new scaffold, we further optimized PPAR activity by enlargement of the lipophilic backbone by substituting the 2,3-dimethylphenyl with biphenylic moieties. Such a substitution pattern had only minor impact on PPAR gamma activity but further increased PPAR alpha activity leading to nanomolar activities. Supporting docking studies proposed that the (R)-enantiomer should fit the PPAR alpha ligand-binding pocket better and thus be more active than the (S)-enantiomer. Single enantiomers of selected active analogues were then prepared by enantio-selective synthesis and enantio-selective preparative HPLC respectively. Biological data for the distinct enantiomers fully corroborated the docking experiments and substantiate a stereochemical impact on PPAR activation.
The subject of this paper is to present molecular descriptors representing the electronegativity of OMO (occupied molecular orbital) and UMO (unoccupied molecular orbital) quantum molecular states that could be used to obtain information on the mechanism of electron transfer between metal compound and biological receptor. The molecular descriptors that were used suggested that the (s,p) and (d(N)) metal ions have different mechanisms of interaction with the receptor. This result explains why the correlation activity-descriptor is rather poor or practically does not exist when all metal ions are analyzed together (irrespective of their valence shell). Since this interaction is the last event of the long chain of processes of the metal ion up to the biological target, such molecular descriptors could be used together with other descriptors in QSAR models for prediction of biological activity.
Choosing a set of molecular descriptors (features) that is most relevant to a given biological response variable is a very important problem in QSAR that has not be solved in an optimal robust way. It is an interesting and important class of mathematical problems, where the number of variables greatly outweighs the number of observations (grossly underdetermined systems). We have used two Bayesian approaches to carry out this task using a suite of QSAR data sets. We employed a specialized sparse Bayesian feature reduction method based on an EM algorithm with a Laplacian prior to select a small set of the most relevant descriptors for modeling the response variables from a much larger pool of possibilities. Having chosen the optimum descriptors in a supervised manner, we used a Bayesian regularized neural network to carry out nonlinear regression and derive robust parsimonious QSAR models for five drug data sets. Models were validated using independent test sets, and results compared with other contemporary descriptor selection methods. Issues around validating small QSAR data sets were also discussed in detail. The sparse feature selection algorithm proved to be an excellent, robust method for selecting descriptors for QSAR models, as it is supervised (descriptors chosen in a context-dependent manner), parsimonious (models not overly complex), and inherently interpretable. Coupled to a robust parsimonious nonlinear modeling method such as the Bayesian regularized neural net, the combination provides a means of optimally modeling the data, and allowing interpretation of the model in terms of the most relevant descriptors.
Multivariate Image Analysis Applied to Quantitative Structure-Activity Relationships (MIA-QSAR) has been recently implemented as a method to model and predict biological activities of drug-like compounds. This method is based on the treatment of 2-D chemical structures, which can be built using specific packages for chemical drawing. These chemical structures correlate with the corresponding bioactivities through descriptors, which are pixels (binaries) of the 2-D images; the variable moiety of chemical structures (substituent groups) explains the variance in the bioactivities column vector of a series of compounds. Thus, the way in which chemical structures are drawn (font type and size, representation of chemical groups, format in which images are saved) should influence the results of prediction. This work reports the statistics of prediction for a case study, a series of anti-HIV compounds, and reveals that the results of prediction is independent of the way in which molecules are drawn.
Distributed grid technologies are gradually realizing their potential to provide innovative infrastructures for complex scientific and industrial applications in the field of computational chemistry and related application areas. The current paper gives examples of distributed solutions for docking and virtual screening applications in moving the computational paradigms towards collaborative research and grid computing environments. The Chemomentum collaborative computing environment including both hardware and software infrastructure is described. Examples of applications are given i) for docking and virtual screening on multiprotein and multilibrary cases for H5N1 avian influenza and HIV-1 viruses; and ii) QSAR model building related to HIV-1 protease activity and aquatic toxicity.