BACKGROUND:Physicians and patients frequently overestimate likelihood of survival after in-hospital cardiopulmonary resuscitation. Discussions and decisions around resuscitation after in-hospital cardiopulmonary arrest often take place without adequate or accurate information.METHODS:We conducted a retrospective chart review of 470 instances of resuscitation after in-hospital cardiopulmonary arrest. Individuals were randomly assigned to a derivation cohort and a validation cohort. Logistic Regression and Linear Discriminant Analysis were used to perform multivariate analysis of the data. The resultant best performing rule was converted to a weighted integer tool, and thresholds of survival and nonsurvival were determined with an attempt to optimize sensitivity and specificity for survival.RESULTS:A 10-feature rule, using thresholds for survival and nonsurvival, was created; the sensitivity of the rule on the validation cohort was 42.7% and specificity was 82.4%. In the Dartmouth Score (DS), the features of age (greater than 70 years of age), history of cancer, previous cardiovascular accident, and presence of coma, hypotension, abnormal PaO2, and abnormal bicarbonate were identified as the best predictors of nonsurvival. Angina, dementia, and chronic respiratory insufficiency were selected as protective features.CONCLUSIONS:Utilizing information easily obtainable on admission, our clinical prediction tool, the DS, provides physicians individualized information about their patients' probability of survival after in-hospital cardiopulmonary arrest. The DS may become a useful addition to medical expertise and clinical judgment in evaluating and communicating an individual's probability of survival after in-hospital cardiopulmonary arrest after it is validated by other cohorts.
We have developed a suite of protein redesign algorithms that improves realistic in silico modeling of proteins. These algorithms are based on three characteristics that make them unique: (1) improved flexibility of the protein backbone, protein side-chains, and ligand to accurately capture the conformational changes that are induced by mutations to the protein sequence; (2) modeling of proteins and ligands as ensembles of low-energy structures to better approximate binding affinity; and (3) a globally optimal protein design search, guaranteeing that the computational predictions are optimal with respect to the input model. Here, we illustrate the importance of these three characteristics. We then describe OSPREY, a protein redesign suite that implements our protein design algorithms. OSPREY has been used prospectively, with experimental validation, in several biomedically relevant settings. We show in detail how OSPREY has been used to predict resistance mutations and explain why improved flexibility, ensembles, and provability are essential for this application.OSPREY is free and open source under a Lesser GPL license. The latest version is OSPREY 2.0. The program, user manual, and source code are available at www.cs.duke.edu/donaldlab/software.php.osprey@cs.duke.edu.
We have developed a suite of protein redesign algorithms that improves realistic in silico modeling of proteins. These algorithms are based on three characteristics that make them unique: (1) improved flexibility of the protein backbone, protein side-chains, and ligand to accurately capture the conformational changes that are induced by mutations to the protein sequence; (2) modeling of proteins and ligands as ensembles of low-energy structures to better approximate binding affinity; and (3) a globally optimal protein design search, guaranteeing that the computational predictions are optimal with respect to the input model. Here, we illustrate the importance of these three characteristics. We then describe OSPREY, a protein redesign suite that implements our protein design algorithms. OSPREY has been used prospectively, with experimental validation, in several biomedically relevant settings. We show in detail how OSPREY has been used to predict resistance mutations and explain why improved flexibility, ensembles, and provability are essential for this application.
Active site mutations that disrupt drug binding are an important mechanism of drug resistance. Computational methods capable of predicting resistance a priori are poised to become extremely useful tools in the fields of drug discovery and treatment design. In this paper, we describe an approach to predicting drug resistance on the basis of Dead-End Elimination and MM-PBSA that requires no prior knowledge of resistance. Our method utilizes a two-pass search to identify mutations that impair drug binding while maintaining affinity for the native substrate. We use our method to probe resistance in four drug-target systems: isoniazid-enoyl-ACP reductase (tuberculosis), ritonavir-HIV protease (HIV), methotrexate-dihydrofolate reductase (breast cancer and leukemia), and gleevec-ABL kinase (leukemia). We validate our model using clinically known resistance mutations for all four test systems. In all cases, the model correctly predicts the majority of known resistance mutations.
Molecular docking is a computational tool commonly applied in drug discovery projects and fundamental biological studies of protein-ligand interactions. Traditionally, molecular docking is used to address one of three following questions: (i) given a ligand molecule and a protein receptor, predict the binding mode (pose) of the ligand within the context of a receptor, (ii) screen a collection of small-molecules against a receptor and rank ligands by their likelihood of being active, and (iii) given a ligand molecule and a target receptor, predict the binding affinity of the two. Here, we focus on the first two questions, namely ranking and pose prediction. Currently, state-of-the-art docking algorithms predict poses within 2Å of the native pose in a rate lower than ∼60% and in many cases, below 40%. In ranking, their ability to identify active ligands is inconsistent and generally suffers from high false-positive rate. In this thesis we present novel algorithms to enhance the ability of molecular docking to address these two questions. These algorithms do not substitute traditional docking but rather being applied on top of them to provide synergistic effect. Our algorithms improve pose predictions by 0.5-1.0Å and ranking order for 23% of the targets in gold-standard benchmarks. As importantly, the algorithms improve the consistence of the posing and ranking predictions over diverse sets of targets and screening libraries. In addition to the posing and ranking, we present the pharmacophore concept. A pharmacophore is an ensemble of physiochemical descriptors associated with a biological target that elucidates common interaction patterns of ligands with that target. We introduce a novel pharmacophore inference algorithm and demonstrate its utilization in molecular docking. This thesis is outlined as follow. First we introduce the molecular docking approach for pose prediction and ranking. Second, we discuss the pharmacophore concept and present algorithms for pharmacophore inference. Third, we demonstrate the utilization of pharmacophores for pose prediction by re-scoring candidate poses generated by docking algorithms. Finally, we present algorithms to improve ranking by reducing bias in scoring functions employed by docking algorithms.
In this article we describe a computational method that automatically generates chemically relevant compound ideas from an initial molecule, closely integrated with in silico models, and a probabilistic scoring algorithm to highlight the compound ideas most likely to satisfy a user-defined profile of required properties. The new compound ideas are generated using medicinal chemistry 'transformation rules' taken from examples in the literature. We demonstrate that the set of 206 transformations employed is generally applicable, produces a wide range of new compounds, and is representative of the types of modifications previously made to move from lead-like to drug-like compounds. Furthermore, we show that more than 94% of the compounds generated by transformation of typical drug-like molecules are acceptable to experienced medicinal chemists. Finally, we illustrate an application of our approach to the lead that ultimately led to the discovery of duloxetine, a marketed serotonin reuptake inhibitor.
Drug discovery research often relies on the use of virtual screening via molecular docking to identify active hits in compound libraries. An area for improvement among many state-of-the-art docking methods is the accuracy of the scoring functions used to differentiate active from nonactive ligands. Many contemporary scoring functions are influenced by the physical properties of the docked molecule. This bias can cause molecules with certain physical properties to incorrectly score better than others. Since variation in physical properties is inevitable in large screening libraries, it is desirable to account for this bias. In this paper, we present a method of normalizing docking scores using virtually generated decoy sets with matched physical properties. First, our method generates a set of property-matched decoys for every molecule in the screening library. Each library molecule and its decoy set are docked using a state-of-the-art method, producing a set of raw docking scores. Next, the raw docking score of each library molecule is normalized against the scores of its decoys. The normalized score represents the probability that the raw docking score was drawn from the background distribution of nonactive property-matched decoys. Assuming that the distribution of scores of active molecules differs from the nonactive score distribution, we expect that the score of an active compound will have a low probability of having been drawn from the nonactive score distribution. In addition to the use of decoys in normalizing docking scores, we suggest that decoy sets may be a useful tool to evaluate, improve, or develop scoring functions. We show that by analyzing docking scores of library molecules with respect to the docking scores of their virtually generated property-matched decoys, one can gain insight into the advantages, limitations, and reliability of scoring functions.
Virtual docking algorithms are often evaluated on their ability to separate active ligands from decoy molecules. The current state-of-the-art benchmark, the Directory of Useful Decoys (DUD), minimizes bias by including decoys from a library of synthetically feasible molecules that are physically similar yet chemically dissimilar to the active ligands. We show that by ignoring synthetic feasibility, we can compile a benchmark that is comparable to the DUD and less biased with respect to physical similarity.
Ligand-based active site alignment is a widely adopted technique for the structural analysis of protein–ligand complexes. However, existing tools for ligand alignment treat the ligands as rigid objects even though most biological ligands are flexible. We present LigAlign, an automated system for flexible ligand alignment and analysis. When performing rigid alignments, LigAlign produces results consistent with manually annotated structural motifs. In performing flexible alignments, LigAlign automatically produces biochemically reasonable ligand fragmentations and subsequently identifies conserved structural motifs that are not detected by rigid alignment.
MOTIVATION Electron cryo-microscopy can be used to infer 3D structures of large macromolecules with high resolution, but the large amounts of data captured necessitate the development of appropriate statistical models to describe the data generation process, and to perform structure inference. We present a new method for performing ab initio inference of the 3D structures of macromolecules from single particle electron cryo-microscopy experiments using class average images. RESULTS We demonstrate this algorithm on one phantom, one synthetic dataset and three real (experimental) datasets (ATP synthase, V-type ATPase and GroEL). Structures consistent with the known structures were inferred for all datasets. AVAILABILITY The software and source code for this method is available for download from our website: http://compbio.cs.toronto.edu/cryoem/.
Adverse drug reactions (ADR), also known as side-effects, are complex undesired physiologic phenomena observed secondary to the administration of pharmaceuticals. Several phenomena underlie the emergence of each ADR; however, a dominant factor is the drug's ability to modulate one or more biological pathways. Understanding the biological processes behind the occurrence of ADRs would lead to the development of safer and more effective drugs. At present, no method exists to discover these ADR-pathway associations. In this paper we introduce a computational framework for identifying a subset of these associations based on the assumption that drugs capable of modulating the same pathway may induce similar ADRs. Our model exploits multiple information resources. First, we utilize a publicly available dataset pairing drugs with their observed ADRs. Second, we identify putative protein targets for each drug using the protein structure database and in-silico virtual docking. Third, we label each protein target with its known involvement in one or more biological pathways. Finally, the relationships among these information sources are mined using multiple stages of logistic-regression while controlling for over-fitting and multiple-hypothesis testing. As proof-of-concept, we examined a dataset of 506 ADRs, 730 drugs, and 830 human protein targets. Our method yielded 185 ADR-pathway associations of which 45 were selected to undergo a manual literature review. We found 32 associations to be supported by the scientific literature.
Dead‐end elimination (DEE) has emerged as a powerful structure‐based, conformational search technique enabling computational protein redesign. Given a protein with n mutable residues, the DEE criteria guide the search toward identifying the sequence of amino acids with the global minimum energy conformation (GMEC). This approach does not restrict the number of permitted mutations and allows the identified GMEC to differ from the original sequence in up to n residues. In practice, redesigns containing a large number of mutations are often problematic when taken into the wet‐lab for creation via site‐directed mutagenesis. The large number of point mutations required for the redesigns makes the process difficult, and increases the risk of major unpredicted and undesirable conformational changes. Preselecting a limited subset of mutable residues is not a satisfactory solution because it is unclear how to select this set before the search has been performed. Therefore, the ideal approach is what we define as the κ‐restricted redesign problem in which any κ of the n residues are allowed to mutate. We introduce restricted dead‐end elimination (rDEE) as a solution of choice to efficiently identify the GMEC of the restricted redesign (the κGMEC). Whereas existing approaches require n ‐choose‐κ individual runs to identify the κGMEC, the rDEE criteria can perform the redesign in a single search. We derive a number of extensions to rDEE and present a restricted form of the A* conformation search. We also demonstrate a 10‐fold speed‐up of rDEE over traditional DEE approaches on three different experimental systems. © 2009 Wiley Periodicals, Inc. J Comput Chem, 2010
MOTIVATION:An enabling resource for drug discovery and protein function prediction is a large, accurate and actively maintained collection of protein/small-molecule complex structures. Models of binding are typically constructed from these structural libraries by generalizing the observed interaction patterns. Consequently, the quality of the model is dependent on the quality of the structural library. An ideal library should be non-biased and comprehensive, contain high-resolution structures and be actively maintained.RESULTS:We present a new protein/small-molecule database (the PSMDB) that offers a non-redundant set of holo PDB complexes. The database was designed to allow frequent updates through a fully automated process without manual annotation or filtering. Our method of database construction addresses redundancy at both the protein and the small-molecule level. By efficiently handling structures with covalently bound ligands, we allow our database to include a larger number of structures than previous methods. Multiple versions of the database are available at our web site, including structures of split complexes--the proteins without their binding ligands and the non-covalently bound ligands within their native coordinate frame.AVAILABILITY:http://compbio.cs.toronto.edu/psmdb
Motivation: The ability to predict binding profiles for an arbitrary protein can significantly improve the areas of drug discovery, lead optimization and protein function prediction. At present, there are no successful algorithms capable of predicting binding profiles for novel proteins. Existing methods typically rely on manually curated templates or entire active site comparison. Consequently, they perform best when analyzing proteins sharing significant structural similarity with known proteins (i.e. proteins resulting from divergent evolution). These methods fall short when used to characterize the binding profile of a novel active site or one for which a template is not available. In contrast to previous approaches, our method characterizes the binding preferences of sub-cavities within the active site by exploiting a large set of known protein–ligand complexes. The uniqueness of our approach lies not only in the consideration of sub-cavities, but also in the more complete structural representation of these sub-cavities, their parametrization and the method by which they are compared. By only requiring local structural similarity, we are able to leverage previously unused structural information and perform binding inference for proteins that do not share significant structural similarity with known systems. Results: Our algorithm demonstrates the ability to accurately cluster similar sub-cavities and to predict binding patterns across a diverse set of protein–ligand complexes. When applied to two high-profile drug targets, our algorithm successfully generates a binding profile that is consistent with known inhibitors. The results suggest that our algorithm should be useful in structure-based drug discovery and lead optimization. Contact: izharw@cs.toronto.edu; lilien@cs.toronto.edu
The ability to predict ligand binding modes without the aid of wet-lab experiments may accelerate and reduce the cost of drug discovery research. Despite significant recent progress, virtual screening has not yet eliminated the need for wet-lab experiments. For example, after a lead compound has been identified, the precise binding mode is still typically determined by experimental structural biology. This structural knowledge is then employed to guide lead optimization. We present a step toward improving protein-ligand binding mode prediction for a set of ligands known to interact with a common protein. There is thus an important distinction between this work and traditional virtual screening algorithms. Whereas traditional approaches attempt to identify binding ligands from a large database of available compounds, our approach aims to more accurately predict the binding mode for a set of ligands which are already known to bind the target protein. The approach is based on the hypothesis that each active site contains a set of interaction points which binding ligands tend to exploit. In a more traditional context, these interaction points make up a pharmacophoric map. Our algorithm first performs traditional protein-ligand docking for each known binder. The ranked lists of candidate binding modes are then evaluated to identify a set of poses maximally self-consistent with respect to a pharmacophoric map generated from the same poses. We have extensively demonstrated the application of the algorithm to four protein systems (thrombin, cyclin-dependent kinase 2, dihydrofolate reductase, and HIV-1 protease) and attained predictions with an average RMSD < 2.5 A for all tested systems. This represents a typical improvement of 0.5-1.0 A (up to 25%) RMSD over the naive virtual docking predictions. Our algorithm is independent of the docking method and may significantly improve binding mode prediction of virtual docking experiments.
In this paper, we discuss the adaptation of an open-source single-user, single-display molecular visualization application for use in a multi-display, multi-user environment. Jmol, a popular, open-source Java applet for viewing PDB files, is modified in such a manner that allows synchronized coordinated views of the same molecule to be displayed in a multi-display workspace. Each display in the workspace is driven by a separate PC, and coordinated views are achieved through the passing of RasMol script commands over the network. The environment includes a tabletop display capable of sensing touch-input, two large vertical displays, and a TabletPC. The presentation of large molecules is adapted to best take advantage of the different qualities of each display, and a set of interaction techniques that allow groups working in this environment to better collaborate are also presented.