BACKGROUND:In the rational drug design process, an ensemble of conformations obtained from a molecular dynamics simulation plays a crucial role in docking experiments. Some studies have found that Fully-Flexible Receptor (FFR) models predict realistic binding energy accurately and improve scoring to enhance selectiveness. At the same time, methods have been proposed to reduce the high computational costs involved in considering the explicit flexibility of proteins in receptor-ligand docking. This study introduces a novel method to optimize ensemble docking-based experiments by reducing the size of an InhA FFR model at docking runtime and scaling docking workflow invocations on cloud virtual machines.RESULTS:First, in order to find the most affordable cost-benefit pool of virtual machines, we evaluated the performance of the docking workflow invocations in different configurations of Azure instances. Second, we validated the gains obtained by the proposed method based on the quality of the Reduced Fully-Flexible Receptor (RFFR) models produced using AutoDock4.2. The analyses show that the proposed method reduced the model size by approximately 50% while covering at least 86% of the best docking results from the 74 ligands tested. Third, we tested our novel method using AutoDock Vina, a different docking software, and showed the positive accuracy achieved in the resulting RFFR models. Finally, our results demonstrated that the method proposed optimized ensemble docking experiments and is applicable to different docking software. In addition, it detected new binding modes, which would be unreachable if employing only the rigid structure used to generate the InhA FFR model.CONCLUSIONS:Our results showed that the selective method is a valuable strategy for optimizing ensemble docking-based experiments using different docking software. The RFFR models produced by discarding non-promising snapshots from the original model are accurately shaped for a larger number of ligands, and the elapsed time spent in the ensemble docking experiments are considerably reduced.
Sentiment analysis is an important technique to interpret user opinion on products from text, for example, as shared in social media. Recent approaches using deep learning can accurately extract overall sentiment from large datasets. However, extracting sentiment from specific aspects of a product with small training datasets remains a challenge. The automatic classification of sentiments at aspect-level can provide more detailed feedbacks about product and service opinions avoiding manual verification. In this work, we develop two deep learning approaches to classify sentiment at aspect-level using small datasets.
Data clustering is the machine learning task that aims at arranging data into groups (clusters) of objects according to a similarity criterion. From an optimisation perspective, it is a particular kind of NP-hard grouping problem, thus attracting much attention from the evolutionary computation community. In this paper, we propose a novel data clustering algorithm based on a univariate estimation of distribution algorithm, namely Clus-EDA. It employs a medoid-based representation in which the cluster prototypes necessarily coincide with objects from the dataset. We compare Clus-EDA with both traditional non-evolutionary clustering algorithms such as k-means and hierarchical agglomerative clustering, and also with an evolutionary algorithm for clustering, in artificial and synthetic datasets. Our results show that Clus-EDA often outperforms the baseline algorithms with regard to distinct cluster validity criteria.
Molecular Dynamics simulations of protein receptors are an emergent tool in rational drug discovery. Nevertheless, employing Molecular Dynamics trajectories in virtual screening of large repositories is a very costly procedure, which ultimately may become unfeasible. Data clustering have been applied in this context with the goal of reducing the overall computational cost in order to make this task feasible. In this paper, we develop a novel estimation of distribution algorithm called Clus-EDA for clustering entire trajectories using structural features from the substrate-binding cavity of the protein receptor. This novel approach is capable of reducing the original trajectory to about 4% of its original size whilst keeping all relevant information for the analysis of receptor-ligand binding. The resulting partition generated by the estimation of distribution algorithm is further validated by analyzing the interactions between 20 ligands and a Fully-Flexible Receptor model containing a 20 ns Molecular Dynamics simulation trajectory. Results show that Clus-EDA is capable of outperforming traditional clustering algorithms such as k-means and hierarchical clustering by providing the smallest variance of the free energy of binding within the conformations in each cluster.
Protein receptor conformations, obtained from molecular dynamics (MD) simulations, have become a promising treatment of its explicit flexibility in molecular docking experiments applied to drug discovery and development. However, incorporating the entire ensemble of MD conformations in docking experiments to screen large candidate compound libraries is currently an unfeasible task. Clustering algorithms have been widely used as a means to reduce such ensembles to a manageable size. Most studies investigate different algorithms using pairwise Root-Mean Square Deviation (RMSD) values for all, or part of the MD conformations. Nevertheless, the RMSD only may not be the most appropriate gauge to cluster conformations when the target receptor has a plastic active site, since they are influenced by changes that occur on other parts of the structure. Hence, we have applied two partitioning methods (k-means and k-medoids) and four agglomerative hierarchical methods (Complete linkage, Ward's, Unweighted Pair Group Method and Weighted Pair Group Method) to analyze and compare the quality of partitions between a data set composed of properties from an enzyme receptor substrate-binding cavity and two data sets created using different RMSD approaches. Ensembles of representative MD conformations were generated by selecting a medoid of each group from all partitions analyzed. We investigated the performance of our new method for evaluating binding conformation of drug candidates to the InhA enzyme, which were performed by cross-docking experiments between a 20 ns MD trajectory and 20 different ligands. Statistical analyses showed that the novel ensemble, which is represented by only 0.48% of the MD conformations, was able to reproduce 75% of all dynamic behaviors within the binding cavity for the docking experiments performed. Moreover, this new approach not only outperforms the other two RMSD-clustering solutions, but it also shows to be a promising strategy to distill biologically relevant information from MD trajectories, especially for docking purposes.
Molecular dynamics simulations of protein receptors have become an attractive tool for rational drug discovery. However, the high computational cost of employing molecular dynamics trajectories in virtual screening of large repositories threats the feasibility of this task. Computational intelligence techniques have been applied in this context, with the ultimate goal of reducing the overall computational cost so the task can become feasible. Particularly, clustering algorithms have been widely used as a means to reduce the dimensionality of molecular dynamics trajectories. In this paper, we develop a novel methodology for clustering entire trajectories using structural features from the substrate-binding cavity of the receptor in order to optimize docking experiments on a cloud-based environment. The resulting partition was selected based on three clustering validity criteria, and it was further validated by analyzing the interactions between 20 ligands and a fully flexible receptor (FFR) model containing a 20 ns molecular dynamics simulation trajectory. Our proposed methodology shows that taking into account features of the substrate-binding cavity as input for the k-means algorithm is a promising technique for accurately selecting ensembles of representative structures tailored to a specific ligand.
Molecular docking simulations are commonly used to identify and optimize drug candidates by examining the interactions between the target protein and small chemical ligands. This procedure is computationally expensive, especially when the receptor is treated as an ensemble of molecular dynamic conformations, namely the Fully-Flexible Receptor (FFR) model. An FFR model can vary from thousands to millions of conformations. Handling molecular docking experiments on FFR models with flexible ligands still constitutes a big challenge, since it may take hours, days, or even months to be completely executed for a single ligand. Moreover, thousands of molecular docking results are quite hard to be analyzed by a domain expert, who typically explores results starting with FEB and RMSD values. This paper addresses the high computational demand to exhaustively execute molecular docking simulations on FFR models, as well as the problem of accurately selecting a small set of representative docking results to be analyzed by a domain expert. Our approach is twofold: (1) we make use of the wFReDoW environment to decrease the dimension of the FFR model during docking experiments, trying to maintain the quality in the resulting reduced models, and (2) we perform careful analyses on docking results to select a set of representative candidate poses. Our simulation results show that the proposed method is able to achieve a trade-off between accuracy and computational cost. This is evidenced from the accuracy in wFReDoW results, which contain 96% of the snapshots within the set of the 100 best FEB values when only 67% of snapshots from the FFR model were docked.
A wide range of public ligand databases provides currently dozens of millions ligands to users. Consequently, exaustive in silico virtual screening testing with such a high volume of data is particularly expensive. Because of this, there is a demand for the development of new solutions that can reduce the number of testing ligands on their target receptors. Nevertheless, there is no method to reduce effectively that high number in a manageable amount, thus becoming this issue a major challenge of rational drug design. This article presents a comparative analysis among the main public ligand databases by measuring the quality and variations in the values of the molecular descriptors available in each one. It aims to help the development of new methods based on criteria that reduce the set of promising ligands to be tested.
In Rational Drug Design (RDD) [1], the interaction between receptor and ligands is the fundamental principle. In in-silico molecular docking experiments it is investigated the best bind and conformation of a ligand into a receptor [2]; those ligands that obtained better in-silico results are tested by in-vitro experiments; if results are promising, a new drug can be produced. However, currently there are more than 20 millions of small molecules available in repositories [3]. If considering the receptor flexibility, it is very expensive and time-consuming to test all these ligands with a single target protein. One possible strategy is to search ligand libraries and, based on a given criterion, select the most promising. In this work we are interested in select candidate ligands to inhibit the InhA [4] receptor from M. Tuberculosis.
Duncan Dubugras A. Ruiz合作论文数Faculdade de Informatica - PUCRS3