Motivation: Many experts worldwide have highlighted the potential of RNA molecules as drug targets for the chemotherapeutic treatment of a range of diseases....
This report describes a new set of macromolecular descriptors of relevance toprotein QSAR/QSPR studies, protein’s quadratic indices. These descriptors are calculatedfrom the macromolecular pseudograph’s α-carbon atom adjacency matrix. A study of theprotein stability effects for a complete set of alanine substitutions in Arc repressorillustrates this approach. Quantitative Structure-Stability Relationship (QSSR) modelsallow discriminating between near wild-type stability and reduced-stability A-mutants. Alinear discriminant function gives rise to excellent discrimination between 85.4% (35/41)and 91.67% (11/12) of near wild-type stability/reduced stability mutants in training andtest series, respectively. The model’s overall predictability oscillates from 80.49 until82.93, when n varies from 2 to 10 in leave-n-out cross validation procedures. This valuestabilizes around 80.49% when n was
Stochastic-based descriptors generated by the MARCH-INSIDE methodology are applied to the prediction of several properties related to electronic, hydrophobicity and size dependent parameters. Linear Regression is the statistical technique employed finding these quantitative correlations. The model obtained explained more than the 85% of the experimental variance and are comparable in the quality of their statistical parameters and predictive power with those previously reported for the prediction of isoelectric point but using robust methodologies for variable selection like Genetic Algorithm and Partial Least Squared (PLS) as well as neural network for the quantitative correlation.
The blood-brain barrier permeation has been investigated by using a topological substructural molecular design approach (TOPS-MODE). A linear regression model was developed to predict the in vivo blood-brain partitioning coefficient on a data set of 119 compounds, treated as the logarithm of the blood-brain concentration ratio. The final model explained the 70% of the variance and it was validated through the use of an external validation set (33 compounds of the 119, MAE= 0.33), a leave-one-out crossvalidation (q(2) = 0.65, S-press = 0.43), fivefold full crossvalidation (removing 28 compounds in each cycle, MAE = 33, RMSE = 0.43) and the prediction of +/- values for an external test set (85.7% of good prediction). This methodology evidenced that the hydrophobicity increase the blood-brain barrier permeation, while the polar surface and its interaction with the atomic mass of compounds decrease it; suggesting the capacity of the TOPS-MODE descriptors to estimate brain penetration potential of new drug candidates. Finally, by the present approach, positive and negative substructural contributions to the brain permeation were identified, and their possibilities in the lead generation and optimization processes were evaluated. (C) 2004 Wiley-Liss, Inc. and the American Pharmacists Association.
We have developed a classification function that is capable of discriminating between anticoccidial and nonanticoccidial compounds with different structural patterns. For this purpose, we calculated the Markovian electron delocalization negentropies of several compounds. These molecular descriptors, which act as molecular fingerprints, are derived from an electronegativity-weighted stochastic matrix (1Π). The method attempts to describe the delocalization of electrons with time during the process of molecule formation by considering the 3D environment of the atoms. Accordingly, the entropies of this random process are used as molecular descriptors. The present study involves a stochastic generalization of the original idea described by Kier, which concerned the use of molecular negentropies in QSAR. Linear discriminant analysis allowed us to fit the discriminant function. This function has given rise to a good classification of 82.35% (28 anticoccidials out of 34) and 91.8% of inactive compounds (56/61) in training series. An overall classification of 88.42% (84/95) was achieved. Validation of the model was carried out by means of an external predicting series and this gave a global predictability of 93.1%. Finally, we report the experimental assay (more than 95% of lesion control) of two compounds selected from a large data set through virtual screening. We conclude that the approach described here seems to be a promising 3D-QSAR tool based on the mathematical theory of stochastic processes.
This report describes a new set of macromolecular descriptors of relevance to nucleic acid QSAR/QSPR studies, nucleic acids’ quadratic indices. These descriptors are calculated from the macromolecular graph’s nucleotide adjacency matrix. A study of the interaction of the antibiotic Paromomycin with the packaging region of the RNA present in type-1 HIV illustrates this approach. A linear discriminant function gave rise to excellent discrimination between 90.10% (91/101) and 81.82% (9/11) of interacting/noninteracting sites of nucleotides in training and test set, respectively. The LOO crossvalidation procedure was used to assess the stability and predictability of the model. Using this approach, the classification model has shown a LOO global good classification of 91.09%. In addition, the model’s overall predictability oscillates from 89.11% until 87.13%, when n varies from 2 to 3 in leave-n-out jackknife method. This value stabilizes around 88.12% when n was > 3. On the other hand, a linear regression model predicted the local binding affinity constants [log K (10-4M-1)] between a specific nucleotide and the aforementioned antibiotic. The linear model explains almost 92% of the variance of the experimental log K (R = 0.96 and s = 0.07) and LOO press statistics evidenced its predictive ability (q2 = 0.85 and scv = 0.09). These models also permit the interpretation of the driving forces of the interaction process. In this sense, developed equations involve short-reaching (k < 3), middle-reaching (4 < k < 9) and far-reaching (k = 10 or greater) nucleotide’s quadratic indices. This situation points to electronic and topologic nucleotide’s backbone interactions control of the stability profile of Paromomycin-RNA complexes. Consequently, the present approach represents a novel and rather promising way to chem & bioinformatics research.
This report describes a new set of macromolecular descriptors of relevance to nucleic acid QSAR/QSPR studies, nucleic acids’ quadratic indices. These descriptors are calculated from the macromolecular graph’s nucleotide adjacency matrix. A study of the interaction of the antibiotic Paromomycin with the packaging region of the RNA present in type-1 HIV illustrates this approach. A linear discriminant function gave rise to excellent discrimination between 90.10% (91/101) and 81.82% (9/11) of interacting/noninteracting sites of nucleotides in training and test set, respectively. The LOO crossvalidation procedure was used to assess the stability and predictability of the model. Using this approach, the classification model has shown a LOO global good classification of 91.09%. In addition, the model’s overall predictability oscillates from 89.11% until 87.13%, when n varies from 2 to 3 in leave-n-out jackknife method. This value stabilizes around 88.12% when n was > 3. On the other hand, a linear regression model predicted the local binding affinity constants [log K (10-4M-1)] between a specific nucleotide and the aforementioned antibiotic. The linear model explains almost 92% of the variance of the experimental log K (R = 0.96 and s = 0.07) and LOO press statistics evidenced its predictive ability (q2 = 0.85 and scv = 0.09). These models also permit the interpretation of the driving forces of the interaction process. In this sense, developed equations involve short-reaching (k < 3), middle-reaching (4 < k < 9) and far-reaching (k = 10 or greater) nucleotide’s quadratic indices. This situation points to electronic and topologic nucleotide’s backbone interactions control of the stability profile of Paromomycin-RNA complexes. Consequently, the present approach represents a novel and rather promising way to chem & bioinformatics research.
A simple stochastic approach, designed to model the movement of electrons throughout chemical bonds, is introduced. This model makes use of a Markov matrix to codify useful structural information in QSAR. The self-return probabilities of this matrix throughout time ((SR)pi(k)) are then used as molecular descriptors. Firstly, a calculation of (SR)pi(k) is made for a large series of anticancer and non-anticancer chemicals. Then, k-Means Cluster Analysis allows us to split the data series into clusters and ensure a representative design of training and predicting series. Next, we develop a classification function through Linear Discriminant Analysis (LDA). This QSAR discriminates between anticancer compounds and non-active compounds with a correct global classification of 90.5% in the training series. The model also correctly classified 86.07% of the compounds in the predicting series. This classification function is then used to perform a virtual screening of a combinatorial library of coumarins. In this connection, the biological assay of some furocoumarins, selected by virtual screening using the present model, gives good results. In particular, a tetracyclic derivative of 5-methoxypsoralen (5-MOP) has an IC50 against HL-60 tumoral line around 6 to 10 times lower than those for 8-MOP and 5-MOP (reference drugs), respectively. Finally, application of Iso-contribution Zone Analysis (IZA) provides structural interpretation of the biological activity predicted with this QSAR.
The TOPological Sub-Structural MOlecular DEsign (TOPS-MODE) approach has been applied to the study of the soil sorption coefficient of various phenylureas herbicides. A model able to describe more than 93% of the variance in the experimental soil sorption coefficient of 44 phenylureas herbicides was developed with the use of the mentioned approach. In contrast, none of eleven different approaches, including the use of Constitutional, Molecular walk counts, BCUT, Charges indices, 2D autocorrelations, Randic molecular profiles, Geometrical, RDF, 3D Morse, GETAWAY and WHIM descriptors was able to explain more than 91% of the variance in the mentioned property with the same number of descriptors. In addition the TOPS MODE allows a simple interpretation of the model in comparison with others methodologies. In addition, the TOPS-MODE approach permitted to find the contribution of different fragments to the soil sorption coefficients giving to the model a straightforward structural interpretability.
The design of novel anti-HIV compounds has now become a crucial area for scientists working in numerous interrelated fields of science such as molecular biology, medicinal chemistry, mathematical biology, molecular modelling and bioinformatics. In this context, the development of simple but physically meaningful mathematical models to represent the interaction between anti-HIV drugs and their biological targets is of major interest. One such area currently under investigation involves the targets in the HIV-RNA-packaging region. In the work described here, we applied Markov chain theory in an attempt to describe the interaction between the antibiotic paromomycin and the packaging region of the RNA in Type-1 HIV. In this model, a nucleic acid squeezed graph is used. The vertices of the graph represent the nucleotides while the edges are the phosphodiester bonds. A stochastic (Markovian) matrix was subsequently defined on this graph, an operation that codifies the probabilities of interaction between specific nucleotides of HIV-RNA and the antibiotic. The strength of these local interactions can be calculated through an inelastic vibrational model. The successive power of this matrix codifies the probabilities with which the vibrations after drug-RNA interactions vanish along the polynucleotide main chain. The sums of self-return probabilities in the k-vicinity of each nucleotide represent physically meaningful descriptors. A linear discriminant function was developed and gave rise to excellent discrimination in 80.8% of interacting and footprinted nucleotides. The Jackknife method was employed to assess the stability and predictability of the model. On the other hand, a linear regression model predicted the local binding affinity constants between a specific nucleotide and the antibiotic (R(2)=0.91, Q(2)=0.86). These kinds of models could play an important role either in the discovery of new anti-HIV compounds or the study of their mode of action.
A novel method for in silico selection of fluckicidal drugs is introduced. Two QSARs that permit us to discriminate between fasciolicide and non-fasciolicide drugs (the first) and to outline some conclusions about the possible mechanism of action of a chemical (the second) are performed. The first model correctly classified 93.85% of compounds in the training series and 89.5% of the compounds in the predicting one. This model correctly classified 87.7, 93.8, 92.2 and 93.9% of compounds in leave-n-out cross validation procedures when n takes values from 2 to until 6. The model seems to be stable in around 92% of good classification in leave-n-out cross validation analysis when n>6. The second model correctly classified 70% of non-fasciolicide compounds, 85.71% of β-tubulin inhibitors and 100% of proton ionophores in the training set. This model recognizes as proton ionophores 100% of any nitrosalicylanilides in the predicting series. Both models have a low p-level <0.05. Finally, the experimental assay of six organic chemicals by an in vivo test permit us to carry out an assessment of the model with a fairly good 100% agreement between experiment and theoretical prediction.