Reductive amination catalysed by imine reductase (IRED) and reductive aminase (RedAm) enzymes has recently been established as a powerful method for the asymmetric synthesis of chiral amines. While this biocatalytic technology has rapidly progressed from proof of concept to initial industrial applications, its scope and limitations remain to be fully explored. In this work, we report a broad and systematic profiling of reductive amination performance in the sequence space of IREDs and RedAms. This investigation employed an iterative strategy for activity and stereoselectivity screening, guided by chemo- and bioinformatic modelling as well as machine learning. By evaluating the catalytic performance of 175 IREDs against structurally diverse panels of 36 carbonyl compounds and 24 amines, we show that the majority of these enzymes is capable of asymmetric reductive amination at equimolar concentrations of the two substrates (50 mM each). The most effective enzymes identified in this study display sequence characteristics of RedAms, are active on 29–42% of the analysed substrate combinations, and combine high specific activities for the most favourable substrate pair (1.7−27.7 U/mgIRED) with excellent stereoselectivity. Beyond assembling this high-performance enzyme panel, we demonstrate extrapolation from our collected screening data to new substrate combinations by deep learning and the scale-up of selected reactions to a preparative batch size (10 mmol substrate, 200 mL reaction volume), delivering gram amounts of reductive amination products in high yield (63–89%) and optical purity (98% to >99% ee).
Adverse drug reactions caused by molecules binding unintended targets are a major concern in drug discovery. Early identification of such interactions during drug design and development minimizes risks and enhances therapeutic efficacy. While experimental approaches are time-consuming and resource-intensive, in silico virtual screening offers a faster, cost-effective strategy to anticipate off-target effects early in drug design. Here, we present a point cloud-based virtual screening method, designed to perform off-target identification based solely on the chemical composition of the primary drug binding site. By screening multidimensional point clouds representing all potential human binding sites, this approach identifies alternative targets based on shape and physicochemical properties. Notably, it operates independently of the protein's overall structure or sequence. This focus on the binding site broadens the search space to structurally unrelated proteins and enables screening without requiring lead molecule information. Using an experimentally validated test set, we demonstrated the method's ability to identify alternative targets across protein families and predict drug promiscuity, achieving a Top-10 recall of 24% for validated, strongly modulated targets. While ligand-agnostic sequence- (BLASTP) and structure-based (Foldseek) methods show higher overall recall, point cloud-based screening uniquely recovers 13.9% (79.4% of its total recovered hits) of known off-targets with <30% sequence identity at Top-100, which are missed by both BLASTP and Foldseek. Beyond known targets, we found several high-ranking candidates not yet annotated as drug-targets but showing notable cavity similarity despite being structurally unrelated, which we present as testable hypotheses for follow-up.
Human proteins are crucial players in both health and disease. Understanding their molecular landscape is a central topic in biological research. Here, we present an extensive dataset of predicted protein structures for 42,042 distinct human proteins, including splicing variants, derived from the UniProt reference proteome UP000005640. To ensure high quality and comparability, the dataset was generated by combining state-of-the-art modeling-tools AlphaFold 2, OpenFold, and ESMFold, provided within NVIDIA’s BioNeMo platform, as well as homology modeling using Innophore’s CavitomiX platform. Our dataset is offered in both unedited and edited formats for diverse research requirements. The unedited version contains structures as generated by the different prediction methods, whereas the edited version contains refinements, including a dataset of structures without low prediction-confidence regions and structures in complex with predicted ligands based on homologs in the PDB. We are confident that this dataset represents the most comprehensive collection of human protein structures available today, facilitating diverse applications such as structure-based drug design and the prediction of protein function and interactions.
Advancing climate change increases the risk of future infectious disease outbreaks, particularly of zoonotic diseases, by affecting the abundance and spread of viral vectors. Concerningly, there are currently no approved drugs for some relevant diseases, such as the arboviral diseases chikungunya, dengue or zika. The development of novel inhibitors takes 10–15 years to reach the market and faces critical challenges in preclinical and clinical trials, with approximately 30% of trials failing due to side effects. As an early response to emerging infectious diseases, CavitOmiX allows for a rapid computational screening of databases containing 3D point-clouds representing binding sites of approved drugs to identify candidates for off-label use. This process, known as drug repurposing, reduces the time and cost of regulatory approval. Here, we present potential approved drug candidates for off-label use, targeting the ADP-ribose binding site of Alphavirus chikungunya non-structural protein 3. Additionally, we demonstrate a novel in silico drug design approach, considering potential side effects at the earliest stages of drug development. We use a genetic algorithm to iteratively refine potential inhibitors for (i) reduced off-target activity and (ii) improved binding to different viral variants or across related viral species, to provide broad-spectrum and safe antivirals for the future.
Treatment of COVID-19 with a soluble version of ACE2 that binds to SARS-CoV-2 virions before they enter host cells is a promising approach, however it needs to be optimized and adapted to emerging viral variants. The computational workflow presented here consists of molecular dynamics simulations for spike RBD-hACE2 binding affinity assessments of multiple spike RBD/hACE2 variants and a novel convolutional neural network architecture working on pairs of voxelized force-fields for efficient search-space reduction. We identified hACE2-Fc K31W and multi-mutation variants as high-affinity candidates, which we validated in vitro with virus neutralization assays. We evaluated binding affinities of these ACE2 variants with the RBDs of Omicron BA.3, Omicron BA.4/BA.5, and Omicron BA.2.75 in silico. In addition, candidates produced in Nicotiana benthamiana , an expression organism for potential large-scale production, showed a 4.6-fold reduction in half-maximal inhibitory concentration (IC 50 ) compared with the same variant produced in CHO cells and an almost six-fold IC 50 reduction compared with wild-type hACE2-Fc.
In this work, we present DrugSolver CavitomiX, a novel computational pipeline for drug repurposing and identifying ligands and inhibitors of target enzymes. The pipeline is based on cavity point clouds representing physico-chemical properties of the cavity induced solely by the protein. To test the pipeline’s ability to identify inhibitors, we chose enzymes essential for SARS-CoV-2 replication as a test system. The active-site cavities of the viral enzymes main protease (M pro ) and papain-like protease (Pl pro ), as well as of the human transmembrane serine protease 2 (TMPRSS2), were selected as target cavities. Using active-site point-cloud comparisons, it was possible to identify two compounds—flufenamic acid and fusidic acid—which show strong inhibition of viral replication. The complexes from which fusidic acid and flufenamic acid were derived would not have been identified using classical sequence- and structure-based methods as they show very little structural (TM-score: 0.1 and 0.09, respectively) and very low sequence (~ 5%) identity to M pro and TMPRSS2, respectively. Furthermore, a cavity-based off-target screening was performed using acetylcholinesterase (AChE) as an example. Using cavity comparisons, the human carboxylesterase was successfully identified, which is a described off-target for AChE inhibitors.
ABSTRACT The monkeypox virus (MPX) belongs to the Orthopoxvirus genus of the Poxviridae family, is endemic in parts of Africa and causes a disease in humans similar to smallpox. The most recent outbreak of MPX is already affecting 110 countries, with 86,956 confirmed cases since May 2022 and has consequently become a focus of interest. In particular, a molecular understanding of the virus is essential to study infection processes and pathogen-host interactions, predict tropism changes, or guide drug development and drug discovery as well as vaccine development or vaccine adaptation at a very early stage. Herein, we present a study of the structural proteome of the currently circulating MPX: Our consensus analysis of 3,713 genome sequences sampled within a year after the outbreak revealed 10,580 characteristic candidate open reading frames (ORFs). A search in the non-redundant protein database reduced the number of suspected ORFs to 1,079, of which 210 are representative proteins in typical MPX reference genomes. This should serve as a collection of putative proteins within the currently spreading MPX, a compound of information that could support timely drug discovery, mutational analyses, and vaccine development. We, herein, present the so far most comprehensive structural proteome by providing atomistic 3D models of 210 proteins, generated with three state-of-the-art structure prediction methods, including a mutational analysis of the proteome, with a particular focus on the drug-binding sites of tecovirimat and brincidofovir. IMPORTANCE The 2022 outbreak of the monkeypox virus already involves, by April 2023, 110 countries with 86,956 confirmed cases and 119 deaths. Understanding an emerging disease on a molecular level is essential to study infection processes and eventually guide drug discovery at an early stage. To support this, we provide the so far most comprehensive structural proteome of the monkeypox virus, which includes 210 structural models, each computed with three state-of-the-art structure prediction methods. Instead of building on a single-genome sequence, we generated our models from a consensus of 3,713 high-quality genome sequences sampled from patients within 1 year of the outbreak. Therefore, we present an average structural proteome of the currently isolated viruses, including mutational analyses with a special focus on drug-binding sites. Continuing dynamic mutation monitoring within the structural proteome presented here is essential to timely predict possible physiological changes in the evolving virus.
The combination of gold(I) and enzyme catalysis has provided access to a series of nor(pseudo)ephedrine derivatives in a regio- and stereoselective manner. The approach involves developing IPrAuNTf2-catalyzed hydration of 1-phenylprop-2-yn-1-yl acetate or N-(1-phenylprop-2-yn-1-yl)acetamide, followed by (dynamic) asymmetric biotransamination or bioreduction of the corresponding keto ester or keto amide intermediates. Enzyme actions were completely selective towards the modification of the methyl ketones in a highly stereoselective manner, allowing the synthesis of enantio- and diastereomerically enriched products using either racemic or optically active starting materials. Thus, a series of amino alcohol, diol, and diamine derivatives were produced from propargyl esters or amides (57 to 86% isolated yield), the biocatalyst of choice determining the (stereo)selectivity of the overall cascade process (70-99% diastereomeric excess and > 98% enantiomeric excess), and providing access to nor(pseudo)ephedrine compounds in a straightforward manner.
The COVID-19 pandemic continues to pose a substantial threat to human lives and is likely to do so for years to come. Despite the availability of vaccines, searching for efficient small-molecule drugs that are widely available, including in low- and middle-income countries, is an ongoing challenge. In this work, we report the results of a community effort, the “Billion molecules against Covid-19 challenge”, to identify small-molecule inhibitors against SARS-CoV-2 or relevant human receptors. Participating teams used a wide variety of computational methods to screen a minimum of 1 billion virtual molecules against 6 protein targets. Overall, 31 teams participated, and they suggested a total of 639,024 potentially active molecules, which were subsequently ranked to find ‘consensus compounds’. The organizing team coordinated with various contract research organizations (CROs) and collaborating institutions to synthesize and test 878 compounds for activity against proteases (Nsp5, Nsp3, TMPRSS2), nucleocapsid N, RdRP (Nsp12 domain), and (alpha) spike protein S. Overall, 27 potential inhibitors were experimentally confirmed by binding-, cleavage-, and/or viral suppression assays and are presented here. All results are freely available and can be taken further downstream without IP restrictions. Overall, we show the effectiveness of computational techniques, community efforts, and communication across research fields (i.e., protein expression and crystallography, in silico modeling, synthesis and biological assays) to accelerate the early phases of drug discovery.
The COVID‐19 pandemic continues to pose a substantial threat to human lives and is likely to do so for years to come. Despite the availability of vaccines, searching for efficient small‐molecule drugs that are widely available, including in low‐ and middle‐income countries, is an ongoing challenge. In this work, we report the results of an open science community effort, the “Billion molecules against COVID‐19 challenge”, to identify small‐molecule inhibitors against SARS‐CoV‐2 or relevant human receptors. Participating teams used a wide variety of computational methods to screen a minimum of 1 billion virtual molecules against 6 protein targets. Overall, 31 teams participated, and they suggested a total of 639,024 molecules, which were subsequently ranked to find ‘consensus compounds’. The organizing team coordinated with various contract research organizations (CROs) and collaborating institutions to synthesize and test 878 compounds for biological activity against proteases (Nsp5, Nsp3, TMPRSS2), nucleocapsid N, RdRP (only the Nsp12 domain), and (alpha) spike protein S. Overall, 27 compounds with weak inhibition/binding were experimentally identified by binding‐, cleavage‐, and/or viral suppression assays and are presented here. Open science approaches such as the one presented here contribute to the knowledge base of future drug discovery efforts in finding better SARS‐CoV‐2 treatments.
The current COVID-19 pandemic poses a challenge to medical professionals and the general public alike. In addition to vaccination programs and nontherapeutic measures being employed worldwide to encounter SARS-CoV-2, great efforts have been made towards drug development and evaluation. In particular, the main protease (M pro ) makes an attractive drug target due to its high level characterization and relatively little similarity to host proteases. Essentially, antiviral strategies are vulnerable to the effects of viral mutation and an early detection of arising resistances supports a timely counteraction in drug development and deployment. Here we show a significant recent event of mutational dynamics in M pro . Although the protease has a priori been expected to be relatively conserved, we report a remarkable increase in mutational variability in an eight-residue long consecutive region near the active site since December 2021. The location of this event in close proximity to an antiviral-drug binding site may suggest the onset of the development of antiviral resistance. Our findings emphasize the importance of monitoring the mutational dynamics of M pro together with possible consequences arising from amino-acid exchanges emerging in regions critical with regard to the susceptibility of the virus to antivirals targeting the protease.
To date, more than 263 million people have been infected with SARS-CoV-2 during the COVID-19 pandemic. In many countries, the global spread occurred in multiple pandemic waves characterized by the emergence of new SARS-CoV-2 variants. Here we report a sequence and structural-bioinformatics analysis to estimate the effects of amino acid substitutions on the affinity of the SARS-CoV-2 spike receptor binding domain (RBD) to the human receptor hACE2. This is done through qualitative electrostatics and hydrophobicity analysis as well as molecular dynamics simulations used to develop a high-precision empirical scoring function (ESF) closely related to the linear interaction energy method and calibrated on a large set of experimental binding energies. For the latest variant of concern (VOC), B.1.1.529 Omicron, our Halo difference point cloud studies reveal the largest impact on the RBD binding interface compared to all other VOC. Moreover, according to our ESF model, Omicron achieves a much higher ACE2 binding affinity than the wild type and, in particular, the highest among all VOCs except Alpha and thus requires special attention and monitoring.
Introduction The current coronavirus pandemic is being combated worldwide by nontherapeutic measures and massive vaccination programs. Nevertheless, therapeutic options such as severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) main-protease (Mpro) inhibitors are essential due to the ongoing evolution toward escape from natural or induced immunity. While antiviral strategies are vulnerable to the effects of viral mutation, the relatively conserved Mpro makes an attractive drug target: Nirmatrelvir, an antiviral targeting its active site, has been authorized for conditional or emergency use in several countries since December 2021, and a number of other inhibitors are under clinical evaluation. We analyzed recent SARS-CoV-2 genomic data, since early detection of potential resistances supports a timely counteraction in drug development and deployment, and discovered accelerated mutational dynamics of Mpro since early December 2021. Methods We performed a comparative analysis of 10.5 million SARS-CoV-2 genome sequences available by June 2022 at GISAID to the NCBI reference genome sequence NC_045512.2. Amino-acid exchanges within high-quality regions in 69,878 unique Mpro sequences were identified and time- and in-depth sequence analyses including a structural representation of mutational dynamics were performed using in-house software. Results The analysis showed a significant recent event of mutational dynamics in Mpro. We report a remarkable increase in mutational variability in an eight-residue long consecutive region (R188-G195) near the active site since December 2021. Discussion The increased mutational variability in close proximity to an antiviral-drug binding site as described herein may suggest the onset of the development of antiviral resistance. This emerging diversity urgently needs to be further monitored and considered in ongoing drug development and lead optimization.
This work describes the biocatalytic amidation of l-proline with ammonia, resulting in a process with optimized atom efficiency giving prolinamide in an optically pure form (ee >99%). Detailed enzyme and reaction engineering studies are provided.
Abstract The monkeypox virus (MPX) belongs to the orthopoxvirus genus of the Poxviridae family, is endemic in parts of Africa, and causes a disease in humans similar to smallpox. The most recent outbreak of MPX in 2022 is already affecting 19 countries on different continents and has consequently become a focus of interest. In particular, a molecular understanding of the virus is essential to study infection processes and pathogen-host interactions, predict tropism changes, or guide drug development and discovery as well as vaccine development or adaptation at a very early stage. Herein we present a study of the structural genome of the currently emerging MPX virus: our analysis revealed 10,043 characteristic candidate open reading frames (ORFs), and a subsequent BLAST search of the non-redundant protein database and PDB reduced the number of suspected ORFs to 925 and 123 protein sequences, respectively. Finally, we provide the 3D structures of these 123 protein sequences, which were predicted by homology modeling and are available for download.
Since the worldwide outbreak of the infectious disease COVID-19, several studies have been published to understand the structural mechanism of the novel coronavirus SARS-CoV-2. During the infection process, the SARS-CoV-2 spike (S) protein plays a crucial role in the receptor recognition and cell membrane fusion process by interacting with the human angiotensin-converting enzyme 2 (hACE2) receptor. However, new variants of these spike proteins emerge as the virus passes through the disease reservoir. This poses a major challenge for designing a potent antigen for an effective immune response against the spike protein. Through a normal mode analysis (NMA) we identified the highly flexible region in the receptor binding domain (RBD) of SARS-CoV-2, starting from residue 475 up to residue 485. Structurally, the position S477 shows the highest flexibility among them. At the same time, S477 is hitherto the most frequently exchanged amino acid residue in the RBDs of SARS-CoV-2 mutants. Therefore, using MD simulations, we have investigated the role of S477 and its two frequent mutations (S477G and S477N) at the RBD during the binding to hACE2. We found that the amino acid exchanges S477G and S477N strengthen the binding of the SARS-COV-2 spike with the hACE2 receptor.
Levetiracetam is an active pharmaceutical ingredient widely used to treat epilepsy.
Ascorbate oxidases are an enzyme group that has not been explored to a large extent. So far, mainly ascorbate oxidases from plants and only a few from fungi have been described. Although ascorbate oxidases belong to the well-studied enzyme family of multi-copper oxidases, their function is still unclear. In this study, Af_AO1, an enzyme from the fungus Aspergillus flavus, was characterized. Sequence analyses and copper content determination demonstrated Af_AO1 to belong to the multi-copper oxidase family. Biochemical characterization and 3D-modeling revealed a similarity to ascorbate oxidases, but also to laccases. Af_AO1 had a 10-fold higher affinity to ascorbic acid (KM = 0.16 ± 0.03 mM) than to ABTS (KM = 1.89 ± 0.12 mM). Furthermore, the best fitting 3D-model was based on the ascorbate oxidase from Cucurbita pepo var. melopepo. The laccase-like activity of Af_AO1 on ABTS (Vmax = 11.56 ± 0.15 µM/min/mg) was, however, not negligible. On the other hand, other typical laccase substrates, such as syringaldezine and guaiacol, were not oxidized by Af_AO1. According to the biochemical and structural characterization, Af_AO1 was classified as ascorbate oxidase with unusual, laccase-like activity.
Molecular dynamics simulations of comparative model of novel coronavirus 2019-nCoV protease Mpro in complex with 6 different conformations based on the Catalophore point-cloud alignment and (re)- docking of lopinavir into the 2019ncov virus protease model. The docking experiment produced 8 clusters of possible conformations, we chose 6 out of 8 conformers and ran an all-atom 300 ps MD at 310 K (=36.85°C).-The two images in the main folder refers to the docked structures before MD simulation.-The file all_centroids.pse contains the frames representing the centroid of the subsequent MD simulation for each docking cluster.-Each archive contains the centroid in PDB format, the starting frame of the simulation in GRO format and the compressed trajectory in XTC format. In the directory "other_files" there are other data generated during the simulation, i.e. heatmap representing the contact frequency between the ligand atoms and the ones belonging to the homology model.