Analysis of observed protein sequences across all species within the UniProtKB/Swiss-Prot data set reveals CQWW as the shortest absent stretch of amino acids. While DNA can be found encoding the CQWW sequence, it has never been observed to be translated or included in manually curated sets of proteins, existing only in predicted, tentative sequences and in a single mature antibody sequence. We have synthesized this "nullomer" peptide, along with 13 derivatives, reversed, truncated, stereoisomers, and alanine-scanning peptides, conjugated to polyarginine stretches to increase cellular uptake. We observed their impact against a healthy neuronal line and six patient-derived glioblastoma cell lines spanning three clinical subtypes. Results reveal IC50 values averaging 4.9 μM for inhibition of cell survival across tested oncogenic cell lines. High-content phenotypic analysis of cellular features and reverse-phase protein arrays failed to discern a clear mode of action for the nullomer peptide but suggests mitochondrial impairment through the inhibition of GSK3 and isoforms, supported by observations of reduced mitochondrial stain intensities. With a recent increase in interest in nullomer peptides, we see the results in this study as a starting point for further investigation into this potentially therapeutic peptide class.
Protein-protein interactions (PPIs) are some of the most challenging target classes in drug discovery. Highly sensitive detection techniques are required for the identification of chemical modulators of PPIs. Here, we introduce PPI confocal nanoscanning (PPI-CONA), a miniaturized, microbead based high-resolution fluorescence imaging assay. We demonstrate the capabilities of PPI-CONA by detecting low affinity ternary complex formation between the human CDC34A ubiquitin-conjugating (E2) enzyme, ubiquitin, and CC0651, a small molecule enhancer of the CDC34A-ubiquitin interaction. We further exemplify PPI-CONA with an E2 enzyme binding study on CC0651 and a CDC34A binding specificity study of a series of CC0651 analogues. Our results indicate that CC0651 is highly selective toward CDC34A. We further demonstrate how PPI-CONA can be applied to screening very low affinity interactions. PPI-CONA holds potential for high-throughput screening for modulators of PPI targets and characterization of their affinity, specificity, and selectivity.
Currently, primaquine is the only malaria transmission-blocking drug recommended by the WHO. Recent efforts have highlighted the importance of discovering new agents that regulate malarial transmission, with particular interest in agents that can be administered in a single low dose, ideally with a discrete and Plasmodium-selective mechanism of action. Here, our team demonstrates an approach to identify malaria transmission-blocking agents through a combination of in vitro screening and in vivo analyses. Using a panel of natural products, our approach identified potent transmission blockers, as illustrated by the discovery of the transmission-blocking efficacy of brusatol. As a member of a large family of biologically active natural products, this discovery provides a critical next step toward developing methods to rapidly identify quassinoids and related agents with valuable pharmacological therapeutic properties.
The Connectivity Map (CMap) is a large publicly available database of cellular transcriptomic responses to chemical and genetic perturbations built using a standardized acquisition protocol known as the L1000 technique. Databases such as CMap provide an exciting opportunity to enrich drug discovery efforts, providing a 'known' phenotypic landscape to explore and enabling the development of state of the art techniques for enhanced information extraction and better informed decisions. Whilst multiple methods for measuring phenotypic similarity and interrogating profiles have been developed, the field is severely lacking standardized benchmarks using appropriate data splitting for training and unbiased evaluation of machine learning methods. To address this, we have developed 'Leak Proof CMap' and exemplified its application to a set of common transcriptomic and generic phenotypic similarity methods along with an exemplar triplet loss-based method. Benchmarking in three critical performance areas (compactness, distinctness, and uniqueness) is conducted using carefully crafted data splits ensuring no similar cell lines or treatments with shared or closely matching responses or mechanisms of action are present in training, validation, or test sets. This enables testing of models with unseen samples akin to exploring treatments with novel modes of action in novel patient derived cell lines. With a carefully crafted benchmark and data splitting regime in place, the tooling now exists to create performant phenotypic similarity methods for use in personalized medicine (novel cell lines) and to better augment high throughput phenotypic screening technologies with the L1000 transcriptomic technology.
The RFamide family of peptides represents an important class of GPCR ligand neuropeptides covering a wide range of biological functions. While many analogues of the highly conserved C-terminal RFamide motif within this peptide class have been synthesized and their functional significance elucidated, additional exploration of the structure activity relationship is of value. We have developed a novel linker for solid phase peptide synthesis (SPPS) which is able to anchor amine functionalised compounds for further elaboration. The acid labile benzofuranone based amine (ALBA) linker (5-(3-aminopropylcarbamoyl)-2-[[tert-butyl(diphenyl)silyl]oxymethyl]benzoic acid) is compatible with Fmoc based SPPS and has two cleavage modes. As a proof of concept, the ALBA linker was used to successfully synthesise a novel analogue of Kisspeptin 10, the natural ligand for GPCR54, whereby the natural RFamide motif was replaced with an RFamine. Biological evaluation of the amine-containing analogue revealed that the group is not compatible with receptor activation. image
SUMMARY:Data integration workflows for multiomics data take many forms across academia and industry. Efforts with limited resources often encountered in academia can easily fall short of data integration best practices for processing and combining high-content imaging, proteomics, metabolomics, and other omics data. We present Phenonaut, a Python software package designed to address the data workflow needs of migration, control, integration, and auditability in the application of literature and proprietary techniques for data source and structure agnostic workflow creation. AVAILABILITY AND IMPLEMENTATION:Source code: https://github.com/CarragherLab/phenonaut, Documentation: https://carragherlab.github.io/phenonaut, PyPI package: https://pypi.org/project/phenonaut/.
Within drug discovery, the goal of AI scientists and cheminformaticians is to help identify molecular starting points that will develop into safe and efficacious drugs while reducing costs, time and failure rates. To achieve this goal, it is crucial to represent molecules in a digital format that makes them machine-readable and facilitates the accurate prediction of properties that drive decision-making. Over the years, molecular representations have evolved from intuitive and human-readable formats to bespoke numerical descriptors and fingerprints, and now to learned representations that capture patterns and salient features across vast chemical spaces. Among these, sequence-based and graph-based representations of small molecules have become highly popular. However, each approach has strengths and weaknesses across dimensions such as generality, computational cost, inversibility for generative applications and interpretability, which can be critical in informing practitioners’ decisions. As the drug discovery landscape evolves, opportunities for innovation continue to emerge. These include the creation of molecular representations for high-value, low-data regimes, the distillation of broader biological and chemical knowledge into novel learned representations and the modeling of up-and-coming therapeutic modalities.
The field of high content imaging has steadily evolved and expanded substantially across many industry and academic research institutions since it was first described in the early 1990 ' s. High content imaging refers to the automated acquisition and analysis of microscopic images from a variety of biological sample types. Integration of high content imaging microscopes with multiwell plate handling robotics enables high content imaging to be performed at scale and support medium-to high -throughput screening of pharmacological, genetic and diverse environmental perturbations upon complex biological systems ranging from 2D cell cultures to 3D tissue organoids to small model organisms. In this perspective article the authors provide a collective view on the following key discussion points relevant to the evolution of high content imaging:center dot Evolution and impact of high content imaging: An academic perspective center dot Evolution and impact of high content imaging: An industry perspective center dot Evolution of high content image analysis center dot Evolution of high content data analysis pipelines towards multiparametric and phenotypic profiling applications center dot The role of data integration and multiomics center dot The role and evolution of image data repositories and sharing standards center dot Future of content hardware and software
A simplistic assumption in setting up a competition assay is that a low affinity labeled ligand can be more easily displaced from a target protein than a high affinity ligand, which in turn produces a more sensitive assay. An often-cited paper correctly rallies against this assumption and recommends the use of the highest affinity ligand available for experiments aiming to determine competitive inhibitor affinities. However, we have noted this advice being applied incorrectly to competition-based primary screens where the goal is optimum assay sensitivity, enabling a clear yes/no binding determination for even low affinity interactions. The published advice only applies to secondary, confirmatory assays intended for accurate affinity determination of primary screening hits. We demonstrate that using very high affinity ligands in competition-based primary screening can lead to reduced assay sensitivity and, ultimately, the discarding of potentially valuable active compounds. We build on techniques developed in our PyBindingCurve software for a mechanistic understanding of complex biological interaction systems, developing the “CLAffinity tool” for simulating competition experiments using protein, ligand, and inhibitor concentrations common to drug screening campaigns. CLAffinity reveals optimum labeled ligand affinity ranges based on assay parameters, rather than general rules to optimize assay sensitivity. We provide the open source CLAffinity software toolset to carry out assay simulations and a video summarizing key findings to aid in understanding, along with a simple lookup table allowing identification of optimal dynamic ranges for competition-based primary screens. The application of our freely available software and lookup tables will lead to the consistent creation of more performant competition-based primary screens identifying valuable hit compounds, particularly for difficult targets.
Background: Lower survival rates for many cancer types correlate with changes in nuclear size/scaling in a tumor-type/tissue-specific manner. Hypothesizing that such changes might confer an advantage to tumor cells, we aimed at the identification of commercially available compounds to guide further mechanistic studies. We therefore screened for Food and Drug Administration (FDA)/European Medicines Agency (EMA)-approved compounds that reverse the direction of characteristic tumor nuclear size changes in PC3, HCT116, and H1299 cell lines reflecting, respectively, prostate adenocarcinoma, colonic adenocarcinoma, and small-cell squamous lung cancer. Results: We found distinct, largely nonoverlapping sets of compounds that rectify nuclear size changes for each tumor cell line. Several classes of compounds including, e.g., serotonin uptake inhibitors, cyclo-oxygenase inhibitors, β-adrenergic receptor agonists, and Na+/K+ ATPase inhibitors, displayed coherent nuclear size phenotypes focused on a particular cell line or across cell lines and treatment conditions. Several compounds from classes far afield from current chemotherapy regimens were also identified. Seven nuclear size-rectifying compounds selected for further investigation all inhibited cell migration and/or invasion. Conclusions: Our study provides (a) proof of concept that nuclear size might be a valuable target to reduce cell migration/invasion in cancer treatment and (b) the most thorough collection of tool compounds to date reversing nuclear size changes specific to individual cancer-type cell lines. Although these compounds still need to be tested in primary cancer cells, the cell line-specific nuclear size and migration/invasion responses to particular drug classes suggest that cancer type-specific nuclear size rectifiers may help reduce metastatic spread.
Small molecule lipophilicity is often included in generalized rules for medicinal chemistry. These rules aim to reduce time, effort, costs, and attrition rates in drug discovery, allowing the rejection or prioritization of compounds without the need for synthesis and testing. The availability of high quality, abundant training data for machine learning methods can be a major limiting factor in building effective property predictors. We utilize transfer learning techniques to get around this problem, first learning on a large amount of low accuracy predicted logP values before finally tuning our model using a small, accurate dataset of 244 druglike compounds to create MRlogP, a neural network-based predictor of logP capable of outperforming state of the art freely available logP prediction methods for druglike small molecules. MRlogP achieves an average root mean squared error of 0.988 and 0.715 against druglike molecules from Reaxys and PHYSPROP. We have made the trained neural network predictor and all associated code for descriptor generation freely available. In addition, MRlogP may be used online via a web interface.
Understanding multicomponent binding interactions in protein-ligand, protein-protein and competition systems is essential for fundamental biology and drug discovery. Hand deriving equations quickly becomes unfeasible when the number of components is increased, and direct analytical solutions only exist to a certain complexity. To address this problem and allow easy access to simulation, plotting and parameter fitting to complex systems at equilibrium, we present the Python package PyBindingCurve. We apply this software to explore homodimer and heterodimer formation culminating in the discovery that under certain conditions, homodimers are easier to break with an inhibitor than heterodimers and may also be more readily depleted. This is a potentially valuable and overlooked phenomenon of great importance to drug discovery. PyBindingCurve may be expanded to operate on any equilibrium binding system and allows definition of custom systems using a simple syntax. PyBindingCurve is available under the MIT license at: https://github.com/stevenshave/pybindingcurve as Python source code accompanied by examples and as an easily installable package within the Python Package Index.
Exploration of chemical space around hit, experimental, and known active compounds is an important step in the early stages of drug discovery. In academia, where access to chemical synthesis efforts is restricted in comparison to the pharma-industry, hits from primary screens are typically followed up through purchase and testing of similar compounds, before further funding is sought to begin medicinal chemistry efforts. Rapid exploration of druglike similars and structure–activity relationship profiles can be achieved through our new webservice SimilarityLab. In addition to searching for commercially available molecules similar to a query compound, SimilarityLab also enables the search of compounds with recorded activities, generating consensus counts of activities, which enables target and off-target prediction. In contrast to other online offerings utilizing the USRCAT similarity measure, SimilarityLab’s set of commercially available small molecules is consistently updated, currently containing over 12.7 million unique small molecules, and not relying on published databases which may be many years out of date. This ensures researchers have access to up-to-date chemistries and synthetic processes enabling greater diversity and access to a wider area of commercial chemical space. All source code is available in the SimilarityLab source repository.
The identification of modulators for proteins without assayable biochemical activity remains a challenge in chemical biology. The presented approach adapts a high-throughput fluorescence binding assay and functional chromatography, two protein-resin technologies, enabling the discovery and isolation of fluorescent natural product probes that target proteins independently of biochemical function. The resulting probes also suggest targetable pockets for lead discovery. Using human survivin as a model, we demonstrate this method with the discovery of members of the prodiginine family as fluorescent probes to the cancer target survivin.
Quantitative microdialysis is a traditional biophysical affinity determination technique. In the development of the detailed experimental protocol presented, we used commercially available equipment, rapid equilibrium dialysis (RED) devices (ThermoFisher Scientific), which means that it is open to most laboratories. The target protein and test compound are incubated in a chamber partitioned to allow only small molecules to transition to a larger reservoir chamber, then reversed-phase high performance liquid chromatography (RP-HPLC) or liquid chromatography–mass spectrometry (LC–MS) is used to determine the abundance of compound in each chamber. A higher compound concentration measured in the chamber that contains the target protein indicates binding. As a novel, and differentiating contribution, we present a protocol for mathematical analysis of experimental data. We provide the equations and the software to yield dissociation constants for the test compound-target protein complex up to 0.5 mM KD, and we quantitatively discuss the limitations of affinities in relation to measured compound concentrations.
Protein-protein-interaction networks (PPINs) organize fundamental biological processes, but how oncogenic mutations impact these interactions and their functions at a network-level scale is poorly understood. Here, we analyze how a common oncogenic KRAS mutation (KRASG13D) affects PPIN structure and function of the Epidermal Growth Factor Receptor (EGFR) network in colorectal cancer (CRC) cells. Mapping >6000 PPIs shows that this network is extensively rewired in cells expressing transforming levels of KRASG13D (mtKRAS). The factors driving PPIN rewiring are multifactorial including changes in protein expression and phosphorylation. Mathematical modelling also suggests that the binding dynamics of low and high affinity KRAS interactors contribute to rewiring. PPIN rewiring substantially alters the composition of protein complexes, signal flow, transcriptional regulation, and cellular phenotype. These changes are validated by targeted and global experimental analysis. Importantly, genetic alterations in the most extensively rewired PPIN nodes occur frequently in CRC and are prognostic of poor patient outcomes.
In this study, we apply a battery of molecular similarity techniques to known inhibitors of kynurenine 3-monooxygenase (KMO), querying each against a repository of approved, experimental, nutraceutical, and illicit drugs. Four compounds are assayed against KMO. Subsequently, diclofenac (also known by the trade names Voltaren, Voltarol, Aclonac, and Cataflam) has been confirmed as a human KMO protein binder and inhibitor in cell lysate with low micromolar KD and IC50, respectively, and low millimolar cellular IC50. Hit to drug hopping, as exemplified here for one of the most successful anti-inflammatory medicines ever invented, holds great promise for expansion into new disease areas and highlights the not-yet-fully-exploited potential of drug repurposing.
Background: The ubiquitin-proteasome system (UPS) controls the stability, localization and/or activity of the proteome. However, the identification and characterization of complex individual ubiquitination cascades and their modulators remains a challenge. Here, we report a broadly applicable, multiplexed, miniaturized on-bead technique for real-time monitoring of various ubiquitination-related enzymatic activities. The assay, termed UPS-confocal fluorescence nanoscanning (UPS-CONA), employs a substrate of interest immobilized on a micro-bead and a fluorescently labeled ubiquitin which, upon enzymatic conjugation to the substrate, is quantitatively detected on the bead periphery by confocal imaging. Results: UPS-CONA is suitable for studying individual enzymatic activities, including various E1, E2, and HECT-type E3 enzymes, and for monitoring multi-step reactions within ubiquitination cascades in a single experimental compartment. We demonstrate the power of the UPS-CONA technique by simultaneously following ubiquitin transfer from Ube1 through Ube2L3 to E6AP. We applied this multi-step setup to investigate the selectivity of five ubiquitination inhibitors reportedly targeting different classes of ubiquitination enzymes. Using UPS-CONA, we have identified a new activity of a small molecule E2 inhibitor, BAY 11-7082, and of a HECT E3 inhibitor, heclin, towards the Ube1 enzyme. Conclusions: As a sensitive, quantitative, flexible, and reagent-efficient method with a straightforward protocol, UPS-CONA constitutes a powerful tool for interrogation of ubiquitination-related enzymatic pathways and their chemical modulators, and is readily scalable for large experiments.
The design of highly diverse phage display libraries is based on assumption that DNA bases are incorporated at similar rates within the randomized sequence. As library complexity increases and expected copy numbers of unique sequences decrease, the exploration of library space becomes sparser and the presence of truly random sequences becomes critical. We present the program PuLSE (Phage Library Sequence Evaluation) as a tool for assessing randomness and therefore diversity of phage display libraries. PuLSE runs on a collection of sequence reads in the fastq file format and generates tables profiling the library in terms of unique DNA sequence counts and positions, translated peptide sequences, and normalized ‘expected’ occurrences from base to residue codon frequencies. The output allows at-a-glance quantitative quality control of a phage library in terms of sequence coverage both at the DNA base and translated protein residue level, which has been missing from toolsets and literature. The open source program PuLSE is available in two formats, a C++ source code package for compilation and integration into existing bioinformatics pipelines and precompiled binaries for ease of use.