The challenges of drug discovery from hit identification to clinical development sometimes involves addressing scaffold hopping issues, in order to optimise molecular biological activity or ADME properties, or mitigate toxicology concerns of a drug candidate. Docking is usually viewed as the method of choice for identification of isofunctional molecules, i. e. highly dissimilar molecules that share common binding modes with a protein target. However, the structure of the protein may not be suitable for docking because of a low resolution, or may even be unknown. This problem is frequently encountered in the case of membrane proteins, although they constitute an important category of the druggable proteome. In such cases, ligand-based approaches offer promise but are often inadequate to handle large-step scaffold hopping, because they usually rely on molecular structure. Therefore, we propose the Interaction Fingerprints Profile (IFPP), a molecular representation that captures molecules binding modes based on docking experiments against a panel of diverse high-quality proteins structures. Evaluation on the LH benchmark demonstrates the interest of IFPP for identification of isofunctional molecules. Nevertheless, computation of IFPPs is expensive, which limits its scalability for screening very large molecular libraries. We propose to overcome this limitation by leveraging Metric Learning approaches, allowing fast estimation of molecules IFPP similarities, thus providing an efficient pre-screening strategy that in applicable to very large molecular libraries. Overall, our results suggest that IFPP provides an interesting and complementary tool alongside existing methods, in order to address challenging scaffold hopping problems effectively in drug discovery.
Drug-target interactions (DTIs) prediction algorithms are used at various stages of the drug discovery process. In this context, specific problems such as deorphanization of a new therapeutic target or target identification of a drug candidate arising from phenotypic screens require large-scale predictions across the protein and molecule spaces. DTI prediction heavily relies on supervised learning algorithms that use known DTIs to learn associations between molecule and protein features, allowing for the prediction of new interactions based on learned patterns. The algorithms must be broadly applicable to enable reliable predictions, even in regions of the protein or molecule spaces where data may be scarce. In this paper, we address two key challenges to fulfill these goals: building large, high-quality training datasets and designing prediction methods that can scale, in order to be trained on such large datasets. First, we introduce LCIdb, a curated, large-sized dataset of DTIs, offering extensive coverage of both the molecule and druggable protein spaces. Notably, LCIdb contains a much higher number of molecules than publicly available benchmarks, expanding coverage of the molecule space. Second, we propose Komet (Kronecker Optimized METhod), a DTI prediction pipeline designed for scalability without compromising performance. Komet leverages a three-step framework, incorporating efficient computation choices tailored for large datasets and involving the Nyström approximation. Specifically, Komet employs a Kronecker interaction module for (molecule, protein) pairs, which efficiently captures determinants in DTIs, and whose structure allows for reduced computational complexity and quasi-Newton optimization, ensuring that the model can handle large training sets, without compromising on performance. Our method is implemented in open-source software, leveraging GPU parallel computation for efficiency. We demonstrate the interest of our pipeline on various datasets, showing that Komet displays superior scalability and prediction performance compared to state-of-the-art deep learning approaches. Additionally, we illustrate the generalization properties of Komet by showing its performance on an external dataset, and on the publicly available LH benchmark designed for scaffold hopping problems. Komet is available open source at https://komet.readthedocs.io and all datasets, including LCIdb, can be found at https://zenodo.org/records/10731712.
Identification of novel chemotypes with biological activity similar to a known active molecule is an important challenge in drug discovery called 'scaffold hopping'. Small-, medium-, and large-step scaffold hopping efforts may lead to increasing degrees of chemical structure novelty with respect to the parent compound. In the present paper, we focus on the problem of large-step scaffold hopping. We assembled a high quality and well characterized dataset of scaffold hopping examples comprising pairs of active molecules and including a variety of protein targets. This dataset was used to build a benchmark corresponding to the setting of real-life applications: one active molecule is known, and the second active is searched among a set of decoys chosen in a way to avoid statistical bias. This allowed us to evaluate the performance of computational methods for solving large-step scaffold hopping problems. In particular, we assessed how difficult these problems are, particularly for classical 2D and 3D ligand-based methods. We also showed that a machine-learning chemogenomic algorithm outperforms classical methods and we provided some useful hints for future improvements.
Accurate prediction of binding affinities from protein-ligand atomic coordinates remains a major challenge in early stages of drug discovery. Using modular message passing graph neural networks describing both the ligand and the protein in their free and bound states, we unambiguously evidence that an explicit description of protein-ligand noncovalent interactions does not provide any advantage with respect to ligand or protein descriptors. Simple models, inferring binding affinities of test samples from that of the closest ligands or proteins in the training set, already exhibit good performances, suggesting that memorization largely dominates true learning in the deep neural networks. The current study suggests considering only noncovalent interactions while omitting their protein and ligand atomic environments. Removing all hidden biases probably requires much denser protein-ligand training matrices and a coordinated effort of the drug design community to solve the necessary protein-ligand structures.
C407 is a compound that corrects the Cystic Fibrosis Transmembrane Conductance Regulator (CFTR) protein carrying the p.Phe508del (F508del) mutation. We investigated the corrector effect of c407 and its derivatives on F508del-CFTR protein. Molecular docking and dynamics simulations combined with site-directed mutagenesis suggested that c407 stabilizes the F508del-Nucleotide Binding Domain 1 (NBD1) during the co-translational folding process by occupying the position of the p.Phe1068 side chain located at the fourth intracellular loop (ICL4). After CFTR domains assembly, c407 occupies the position of the missing p.Phe508 side chain. C407 alone or in combination with the F508del-CFTR corrector VX-809, increased CFTR activity in cell lines but not in primary respiratory cells carrying the F508del mutation. A structure-based approach resulted in the synthesis of an extended c407 analog G1, designed to improve the interaction with ICL4. G1 significantly increased CFTR activity and response to VX-809 in primary nasal cells of F508del homozygous patients. Our data demonstrate that in-silico optimized c407 derivative G1 acts by a mechanism different from the reference VX-809 corrector and provide insights into its possible molecular mode of action. These results pave the way for novel strategies aiming to optimize the flawed ICL4–NBD1 interface.
Recent evidence shows that combination of correctors and potentiators, such as the drug ivacaftor (VX-770), can significantly restore the functional expression of mutated Cystic Fibrosis Transmembrane conductance Regulator (CFTR), an anion channel which is mutated in cystic fibrosis (CF). The success of these combinatorial therapies highlights the necessity of identifying a broad panel of specific binding mode modulators, occupying several distinct binding sites at structural level. Here, we identified two small molecules, SBC040 and SBC219, which are two efficient cAMP-independent potentiators, acting at low concentration of forskolin with EC50 close to 1 μM and in a synergic way with the drug VX-770 on several CFTR mutants of classes II and III. Molecular dynamics simulations suggested potential SBC binding sites at the vicinity of ATP-binding sites, distinct from those currently proposed for VX-770, outlining SBC molecules as members of a new family of potentiators.
Understanding the functional consequence of rare cystic fibrosis (CF) mutations is mandatory for the adoption of precision therapeutic approaches for CF. Here we studied the effect of the very rare CF mutation, W361R, on CFTR processing and function. We applied western blot, patch clamp and pharmacological modulators of CFTR to study the maturation and ion transport properties of pEGFP-WT and mutant CFTR constructs, W361R, F508del and L69H-CFTR, expressed in HEK293 cells. Structural analyses were also performed to study the molecular environment of the W361 residue. Western blot showed that W361R-CFTR was not efficiently processed to a mature band C, similar to F508del CFTR, but unlike F508del CFTR, it did exhibit significant transport activity at the cell surface in response to cAMP agonists. Importantly, W361R-CFTR also responded well to CFTR modulators: its maturation defect was efficiently corrected by VX-809 treatment and its channel activity further potentiated by VX-770. Based on these results, we postulate that W361R is a novel class-2 CF mutation that causes abnormal protein maturation which can be corrected by VX-809, and additionally potentiated by VX-770, two FDA-approved small molecules. At the structural level, W361 is located within a class-2 CF mutation hotspot that includes other mutations that induce variable disease severity. Analysis of the 3D structure of CFTR within a lipid environment indicated that W361, together with other mutations located in this hotspot, is at the edge of a groove which stably accommodates lipid acyl chains. We suggest this lipid environment impacts CFTR folding, maturation and response to CFTR modulators.
Severe chronic rhinosinusitis in children should alert clinicians and extensive CFTR genotyping should be performed. We propose that thorough clinical and functional assessment in severe chronic rhinosinusitis is valuable to discover rare mutations which could be treated by CFTR correctors to postpone pulmonary infection.
Praziquantel (PZQ) is the first line drug for the treatment of human Schistosoma spp. worm infections. However, it suffers from low activity towards immature stages of the worm, and its prolonged use induces resistance/tolerance. During the last 40 years, 263 PZQ analogues have been synthesized and tested against Schistosoma spp. worms, but less than 10% of them showed significant activity. Here, we propose a rationalization of the chemical space of the PZQ derivatives by a ligand-based approach. First, we constructed an in-house database with all PZQ derivatives available in the literature. This analysis shows a high heterogeneity in the data. Fortunately, all studies include PZQ as a reference, permitting the classification of compounds into three classes according to their activities. Models involving ligand-based pharmacophore and logistic regression were performed. Five physicochemical parameters were identified as the best to explain the biological activity. In the end, we proposed new PZQ derivatives with modifications at positions 1 and 7, we analysed them with our models, and we observed that they can be more active than the previously synthesized derivatives. The main goal of this work was to conduct the most valuable meta-pharmacometrics/pharmacoinformatics analysis with all Praziquantel medicinal chemistry data available in the literature.
Cryo-electron microscopy (cryo-EM) has recently provided invaluable experimental data about the full-length cystic fibrosis transmembrane conductance regulator (CFTR) 3D structure. However, this experimental information deals with inactive states of the channel, either in an apo, quiescent conformation, in which nucleotide-binding domains (NBDs) are widely separated or in an ATP-bound, yet closed conformation. Here, we show that 3D structure models of the open and closed forms of the channel, now further supported by metadynamics simulations and by comparison with the cryo-EM data, could be used to gain some insights into critical features of the conformational transition toward active CFTR forms. These critical elements lie within membrane-spanning domains but also within NBD1 and the N-terminal extension, in which conformational plasticity is predicted to occur to help the interaction with filamin, one of the CFTR cellular partners.
Molecules correcting the trafficking (correctors) and gating defects (potentiators) of the cystic fibrosis causing mutation c.1521_1523delCTT (p.Phe508del) begin to be a useful treatment for CF patients bearing p.Phe508del. This mutation has been identified in different genetic contexts, alone or in combination with variants in cis. Until now, 21 exonic variants in cis of p.Phe508del have been identified, albeit at a low frequency. The aim of this study was to evaluate their impact on the efficacy of CFTR-directed corrector/potentiator therapy (Orkambi). The analysis by minigene showed that two out of 15 cis variants tested increased exon skipping (c.609C > T and c.2770G > A). Four cis variants were studied functionally in the absence of p.Phe508del, one of which was found to be deleterious for protein maturation c.1399C > T (p.Leu467Phe). In the presence of p.Phe508del, this variant was the only to prevent the response to Orkambi treatment. This study showed that some patients carrying p.Phe508del complex alleles are predicted to poorly respond to corrector/potentiator treatments. Our results underline the importance to validate treatment efficacy in the context of complex alleles.
ABCB4 (MDR3) is an adenosine triphosphate (ATP)‐binding cassette (ABC) transporter expressed at the canalicular membrane of hepatocytes, where it mediates phosphatidylcholine (PC) secretion. Variations in the ABCB4 gene are responsible for several biliary diseases, including progressive familial intrahepatic cholestasis type 3 (PFIC3), a rare disease that can be lethal in the absence of liver transplantation. In this study, we investigated the effect and potential rescue of ABCB4 missense variations that reside in the highly conserved motifs of ABC transporters, involved in ATP binding. Five disease‐causing variations in these motifs have been identified in ABCB4 (G535D, G536R, S1076C, S1176L, and G1178S), three of which are homologous to the gating mutations of cystic fibrosis transmembrane conductance regulator (CFTR or ABCC7; i.e., G551D, S1251N, and G1349D), that were previously shown to be function defective and corrected by ivacaftor (VX‐770; Kalydeco), a clinically approved CFTR potentiator. Three‐dimensional structural modeling predicted that all five ABCB4 variants would disrupt critical interactions in the binding of ATP and thereby impair ATP‐induced nucleotide‐binding domain dimerization and ABCB4 function. This prediction was confirmed by expression in cell models, which showed that the ABCB4 mutants were normally processed and targeted to the plasma membrane, whereas their PC secretion activity was dramatically decreased. As also hypothesized on the basis of molecular modeling, PC secretion activity of the mutants was rescued by the CFTR potentiator, ivacaftor (VX‐770). Conclusion: Disease‐causing variations in the ATP‐binding sites of ABCB4 cause defects in PC secretion, which can be rescued by ivacaftor. These results provide the first experimental evidence that ivacaftor is a potential therapy for selected patients who harbor mutations in the ATP‐binding sites of ABCB4. (Hepatology 2017;65:560‐570)
Development of Cystic Fibrosis Transmembrane conductance Regulator (CFTR) modulators, targeting the root cause of cystic fibrosis (CF), represents a challenge in the era of personalized medicine, as CFTR mutations lead to a variety of phenotypes, which likely require different, specific treatments. CF drug development is also complicated by the need to preserve the right balance between stability and flexibility, required for optimal function of the CFTR protein. In this review, we highlight how structural data can be exploited in this context to understand the molecular mechanisms of disease-associated mutations, to characterize the mechanisms of action of known modulators and to rationalize the search for novel, specific compounds.
The intermediate filament protein keratin 8 (K8) interacts with the nucleotide‐binding domain 1 (NBD1) of the cystic fibrosis (CF) transmembrane regulator (CFTR) with phenylalanine 508 deletion (ΔF508), and this interaction hampers the biogenesis of functional ΔF508‐CFTR and its insertion into the plasma membrane. Interruption of this interaction may constitute a new therapeutic target for CF patients bearing the ΔF508 mutation. Here, we aimed to determine the binding surface between these two proteins, to facilitate the design of the interaction inhibitors. To identify the NBD1 fragments perturbed by the ΔF508 mutation, we used hydrogen–deuterium exchange coupled with mass spectrometry (HDX‐MS) on recombinant wild‐type (wt) NBD1 and ΔF508‐NBD1 of CFTR. We then performed the same analysis in the presence of a peptide from the K8 head domain, and extended this investigation using bioinformatics procedures and surface plasmon resonance, which revealed regions affected by the peptide binding in both wt‐NBD1 and ΔF508‐NBD1. Finally, we performed HDX‐MS analysis of the NBD1 molecules and full‐length K8, revealing hydrogen‐bonding network changes accompanying complex formation. In conclusion, we have localized a region in the head segment of K8 that participates in its binding to NBD1. Our data also confirm the stronger binding of K8 to ΔF508‐NBD1, which is supported by an additional binding site located in the vicinity of the ΔF508 mutation in NBD1.
The cystic fibrosis transmembrane conductance regulator (CFTR) protein is a member of the ATP-binding cassette (ABC) transporter superfamily that functions as an ATP-gated channel. Considerable progress has been made over the last years in the understanding of the molecular basis of the CFTR functions, as well as dysfunctions causing the common genetic disease cystic fibrosis (CF). This review provides a global overview of the theoretical studies that have been performed so far, especially molecular modelling and molecular dynamics (MD) simulations. A special emphasis is placed on the CFTR-specific evolution of an ABC transporter framework towards a channel function, as well as on the understanding of the effects of disease-causing mutations and their specific modulation. This in silico work should help structure-based drug discovery and design, with a view to develop CFTR-specific pharmacotherapeutic approaches for the treatment of CF in the context of precision medicine.