Adverse drug reactions caused by molecules binding unintended targets are a major concern in drug discovery. Early identification of such interactions during drug design and development minimizes risks and enhances therapeutic efficacy. While experimental approaches are time-consuming and resource-intensive, in silico virtual screening offers a faster, cost-effective strategy to anticipate off-target effects early in drug design. Here, we present a point cloud-based virtual screening method, designed to perform off-target identification based solely on the chemical composition of the primary drug binding site. By screening multidimensional point clouds representing all potential human binding sites, this approach identifies alternative targets based on shape and physicochemical properties. Notably, it operates independently of the protein's overall structure or sequence. This focus on the binding site broadens the search space to structurally unrelated proteins and enables screening without requiring lead molecule information. Using an experimentally validated test set, we demonstrated the method's ability to identify alternative targets across protein families and predict drug promiscuity, achieving a Top-10 recall of 24% for validated, strongly modulated targets. While ligand-agnostic sequence- (BLASTP) and structure-based (Foldseek) methods show higher overall recall, point cloud-based screening uniquely recovers 13.9% (79.4% of its total recovered hits) of known off-targets with <30% sequence identity at Top-100, which are missed by both BLASTP and Foldseek. Beyond known targets, we found several high-ranking candidates not yet annotated as drug-targets but showing notable cavity similarity despite being structurally unrelated, which we present as testable hypotheses for follow-up.
Human proteins are crucial players in both health and disease. Understanding their molecular landscape is a central topic in biological research. Here, we present an extensive dataset of predicted protein structures for 42,042 distinct human proteins, including splicing variants, derived from the UniProt reference proteome UP000005640. To ensure high quality and comparability, the dataset was generated by combining state-of-the-art modeling-tools AlphaFold 2, OpenFold, and ESMFold, provided within NVIDIA’s BioNeMo platform, as well as homology modeling using Innophore’s CavitomiX platform. Our dataset is offered in both unedited and edited formats for diverse research requirements. The unedited version contains structures as generated by the different prediction methods, whereas the edited version contains refinements, including a dataset of structures without low prediction-confidence regions and structures in complex with predicted ligands based on homologs in the PDB. We are confident that this dataset represents the most comprehensive collection of human protein structures available today, facilitating diverse applications such as structure-based drug design and the prediction of protein function and interactions.
Advancing climate change increases the risk of future infectious disease outbreaks, particularly of zoonotic diseases, by affecting the abundance and spread of viral vectors. Concerningly, there are currently no approved drugs for some relevant diseases, such as the arboviral diseases chikungunya, dengue or zika. The development of novel inhibitors takes 10–15 years to reach the market and faces critical challenges in preclinical and clinical trials, with approximately 30% of trials failing due to side effects. As an early response to emerging infectious diseases, CavitOmiX allows for a rapid computational screening of databases containing 3D point-clouds representing binding sites of approved drugs to identify candidates for off-label use. This process, known as drug repurposing, reduces the time and cost of regulatory approval. Here, we present potential approved drug candidates for off-label use, targeting the ADP-ribose binding site of Alphavirus chikungunya non-structural protein 3. Additionally, we demonstrate a novel in silico drug design approach, considering potential side effects at the earliest stages of drug development. We use a genetic algorithm to iteratively refine potential inhibitors for (i) reduced off-target activity and (ii) improved binding to different viral variants or across related viral species, to provide broad-spectrum and safe antivirals for the future.
Light-dependent fatty acid photodecarboxylases (FAPs) hold significant potential for biotechnology, due to their capability to produce alka(e)nes directly from the corresponding (un)saturated natural fatty acids requiring light as the only reagent. This study expands the family of FAPs through cavity-based enzyme discovery methods. Thirty enzyme candidates with potential photodecarboxylation activity were identified by matching the cavities of four related template structures against the Protein Data Bank's flavoproteins, a library of proteins identified via the Foldseek Search Server, and homology models of sequences resulting from BLAST. Subsequent docking experiments narrowed this library to ten promising enzymes, which were expressed and assessed in vitro, identifying four photodecarboxylases. Out of these enzymes, the GMC oxidoreductase from Coccomyxa sp. Obi (CoFAP) was characterized in detail, which revealed high activity in the decarboxylation reactions of palmitic acid and octanoic acid and a broad pH tolerance (pH 6.5-9.5).
In this work, we present DrugSolver CavitomiX, a novel computational pipeline for drug repurposing and identifying ligands and inhibitors of target enzymes. The pipeline is based on cavity point clouds representing physico-chemical properties of the cavity induced solely by the protein. To test the pipeline’s ability to identify inhibitors, we chose enzymes essential for SARS-CoV-2 replication as a test system. The active-site cavities of the viral enzymes main protease (M pro ) and papain-like protease (Pl pro ), as well as of the human transmembrane serine protease 2 (TMPRSS2), were selected as target cavities. Using active-site point-cloud comparisons, it was possible to identify two compounds—flufenamic acid and fusidic acid—which show strong inhibition of viral replication. The complexes from which fusidic acid and flufenamic acid were derived would not have been identified using classical sequence- and structure-based methods as they show very little structural (TM-score: 0.1 and 0.09, respectively) and very low sequence (~ 5%) identity to M pro and TMPRSS2, respectively. Furthermore, a cavity-based off-target screening was performed using acetylcholinesterase (AChE) as an example. Using cavity comparisons, the human carboxylesterase was successfully identified, which is a described off-target for AChE inhibitors.
The COVID-19 pandemic continues to pose a substantial threat to human lives and is likely to do so for years to come. Despite the availability of vaccines, searching for efficient small-molecule drugs that are widely available, including in low- and middle-income countries, is an ongoing challenge. In this work, we report the results of a community effort, the “Billion molecules against Covid-19 challenge”, to identify small-molecule inhibitors against SARS-CoV-2 or relevant human receptors. Participating teams used a wide variety of computational methods to screen a minimum of 1 billion virtual molecules against 6 protein targets. Overall, 31 teams participated, and they suggested a total of 639,024 potentially active molecules, which were subsequently ranked to find ‘consensus compounds’. The organizing team coordinated with various contract research organizations (CROs) and collaborating institutions to synthesize and test 878 compounds for activity against proteases (Nsp5, Nsp3, TMPRSS2), nucleocapsid N, RdRP (Nsp12 domain), and (alpha) spike protein S. Overall, 27 potential inhibitors were experimentally confirmed by binding-, cleavage-, and/or viral suppression assays and are presented here. All results are freely available and can be taken further downstream without IP restrictions. Overall, we show the effectiveness of computational techniques, community efforts, and communication across research fields (i.e., protein expression and crystallography, in silico modeling, synthesis and biological assays) to accelerate the early phases of drug discovery.
The COVID‐19 pandemic continues to pose a substantial threat to human lives and is likely to do so for years to come. Despite the availability of vaccines, searching for efficient small‐molecule drugs that are widely available, including in low‐ and middle‐income countries, is an ongoing challenge. In this work, we report the results of an open science community effort, the “Billion molecules against COVID‐19 challenge”, to identify small‐molecule inhibitors against SARS‐CoV‐2 or relevant human receptors. Participating teams used a wide variety of computational methods to screen a minimum of 1 billion virtual molecules against 6 protein targets. Overall, 31 teams participated, and they suggested a total of 639,024 molecules, which were subsequently ranked to find ‘consensus compounds’. The organizing team coordinated with various contract research organizations (CROs) and collaborating institutions to synthesize and test 878 compounds for biological activity against proteases (Nsp5, Nsp3, TMPRSS2), nucleocapsid N, RdRP (only the Nsp12 domain), and (alpha) spike protein S. Overall, 27 compounds with weak inhibition/binding were experimentally identified by binding‐, cleavage‐, and/or viral suppression assays and are presented here. Open science approaches such as the one presented here contribute to the knowledge base of future drug discovery efforts in finding better SARS‐CoV‐2 treatments.
The current COVID-19 pandemic poses a challenge to medical professionals and the general public alike. In addition to vaccination programs and nontherapeutic measures being employed worldwide to encounter SARS-CoV-2, great efforts have been made towards drug development and evaluation. In particular, the main protease (M pro ) makes an attractive drug target due to its high level characterization and relatively little similarity to host proteases. Essentially, antiviral strategies are vulnerable to the effects of viral mutation and an early detection of arising resistances supports a timely counteraction in drug development and deployment. Here we show a significant recent event of mutational dynamics in M pro . Although the protease has a priori been expected to be relatively conserved, we report a remarkable increase in mutational variability in an eight-residue long consecutive region near the active site since December 2021. The location of this event in close proximity to an antiviral-drug binding site may suggest the onset of the development of antiviral resistance. Our findings emphasize the importance of monitoring the mutational dynamics of M pro together with possible consequences arising from amino-acid exchanges emerging in regions critical with regard to the susceptibility of the virus to antivirals targeting the protease.
To date, more than 263 million people have been infected with SARS-CoV-2 during the COVID-19 pandemic. In many countries, the global spread occurred in multiple pandemic waves characterized by the emergence of new SARS-CoV-2 variants. Here we report a sequence and structural-bioinformatics analysis to estimate the effects of amino acid substitutions on the affinity of the SARS-CoV-2 spike receptor binding domain (RBD) to the human receptor hACE2. This is done through qualitative electrostatics and hydrophobicity analysis as well as molecular dynamics simulations used to develop a high-precision empirical scoring function (ESF) closely related to the linear interaction energy method and calibrated on a large set of experimental binding energies. For the latest variant of concern (VOC), B.1.1.529 Omicron, our Halo difference point cloud studies reveal the largest impact on the RBD binding interface compared to all other VOC. Moreover, according to our ESF model, Omicron achieves a much higher ACE2 binding affinity than the wild type and, in particular, the highest among all VOCs except Alpha and thus requires special attention and monitoring.
Introduction The current coronavirus pandemic is being combated worldwide by nontherapeutic measures and massive vaccination programs. Nevertheless, therapeutic options such as severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) main-protease (Mpro) inhibitors are essential due to the ongoing evolution toward escape from natural or induced immunity. While antiviral strategies are vulnerable to the effects of viral mutation, the relatively conserved Mpro makes an attractive drug target: Nirmatrelvir, an antiviral targeting its active site, has been authorized for conditional or emergency use in several countries since December 2021, and a number of other inhibitors are under clinical evaluation. We analyzed recent SARS-CoV-2 genomic data, since early detection of potential resistances supports a timely counteraction in drug development and deployment, and discovered accelerated mutational dynamics of Mpro since early December 2021. Methods We performed a comparative analysis of 10.5 million SARS-CoV-2 genome sequences available by June 2022 at GISAID to the NCBI reference genome sequence NC_045512.2. Amino-acid exchanges within high-quality regions in 69,878 unique Mpro sequences were identified and time- and in-depth sequence analyses including a structural representation of mutational dynamics were performed using in-house software. Results The analysis showed a significant recent event of mutational dynamics in Mpro. We report a remarkable increase in mutational variability in an eight-residue long consecutive region (R188-G195) near the active site since December 2021. Discussion The increased mutational variability in close proximity to an antiviral-drug binding site as described herein may suggest the onset of the development of antiviral resistance. This emerging diversity urgently needs to be further monitored and considered in ongoing drug development and lead optimization.
Monolignol oxidoreductases are members of the berberine bridge enzyme–like (BBE-like) protein family (pfam 08031) that oxidize monolignols to the corresponding aldehydes. They are FAD-dependent enzymes that exhibit the para-cresolmethylhydroxylase-topology, also known as vanillyl oxidase-topology. Recently, we have reported the structural and biochemical characterization of two monolignol oxidoreductases from Arabidopsis thaliana, AtBBE13 and AtBBE15. Now, we have conducted a comprehensive site directed mutagenesis study for AtBBE15, to expand our understanding of the catalytic mechanism of this enzyme class. Based on the kinetic properties of active site variants and molecular dynamics simulations, we propose a refined, structure-guided reaction mechanism for the family of monolignol oxidoreductases. Here, we propose that this reaction is facilitated stepwise by the deprotonation of the allylic alcohol and a subsequent hydride transfer from the Cα-atom of the alkoxide to the flavin. We describe an excessive hydrogen bond network that enables the catalytic mechanism of the enzyme. Within this network Tyr479 and Tyr193 act concertedly as active catalytic bases to facilitate the proton abstraction. Lys436 is indirectly involved in the deprotonation as this residue determines the position of Tyr193 via a cation-π interaction. The enzyme forms a hydrophilic cavity to accommodate the alkoxide intermediate and to stabilize the transition state from the alkoxide to the aldehyde. By means of molecular dynamics simulations, we have identified two different and distinct binding modes for the substrate in the alcohol and alkoxide state. The alcohol interacts with Tyr193 and Tyr479 while Arg292, Gln438 and Tyr193 form an alkoxide binding site to accommodate this intermediate. The pH-dependency of the activity of the active site variants revealed that the integrity of the alkoxide binding site is also crucial for the fine tuning of the pKa of Tyr193 and Tyr479. Sequence alignments showed that key residues for the mechanism are highly conserved, indicating that our proposed mechanism is not only relevant for AtBBE15 but for the majority of BBE-like proteins.
Podophyllotoxin is probably the most prominent representative of lignan natural products. Deoxy-, epi-, and podophyllotoxin, which are all precursors to frequently used chemotherapeutic agents, were prepared by a stereodivergent biotransformation and a biocatalytic kinetic resolution of the corresponding dibenzylbutyrolactones with the same 2-oxoglutarate-dependent dioxygenase. The reaction can be conducted on 2 g scale, and the enzyme allows tailoring of the initial, "natural" structure and thus transforms various non-natural derivatives. Depending on the substitution pattern, the enzyme performs an oxidative C-C bond formation by C-H activation or hydroxylation at the benzylic position prone to ring closure.
AbstractPodophyllotoxin (1) ist einer der bekanntesten Vertreter von natürlich vorkommenden Lignanen. Deoxy‐, epi‐ und Podophyllotoxin, welche alle Vorstufen für häufig verwendete chemotherapeutische Verbindungen darstellen, wurden mithilfe einer stereodivergenten Biotransformation und biokatalytischen kinetischen Racematspaltung der jeweiligen Dibenzylbutyrolactone durch dieselbe 2‐Oxoglutarat‐abhängige Dioxygenase (2‐ODD) hergestellt. Zudem konnte gezeigt werden, dass ein “Upscaling” der Reaktion auf 2 g möglich ist, sowie dass das 2‐ODD‐Enzym Modifikationen des “natürlichen” Substrates toleriert und somit verschiedenste nichtnatürliche Derivate umsetzen kann. Das Enzym führt je nach Substituentenmuster entweder eine oxidative C‐C‐Bindungsbildung durch C‐H‐Aktivierung oder Hydroxylierung an der für die Ringschließung zugänglichen benzylischen Position durch.