After initial triaging using in vitro absorption, distribution, metabolism, and excretion (ADME) assays, pharmacokinetic (PK) studies are the first application of promising drug candidates in living mammals. Pre-clinical PK studies characterize the evolution of the compound’s concentration over time, typically in rodents’ blood or plasma. From this concentration-time (C-t) profiles, PK parameters such as total exposure or maximum concentration can be subsequently derived. An early estimation of compounds’ PK offers the promise of reducing animal studies and cycle times by selecting and designing molecules with increased chances of success at the PK stage. Even though C-t curves are the major readout from a PK study, most machine learning-based prediction efforts have focused on the derived PK parameters instead of C-t profiles, likely due to the lack of approaches to model the underlying ADME mechanisms. Herein, a novel deep learning approach termed DeepCt is proposed for the prediction of C-t curves from the compound structure. Our methodology is based on the prediction of an underlying mechanistic compartmental PK model, which enables further simulations, and predictions of single- and multiple-dose C-t profiles.
In vitro secondary pharmacology assays are an important tool for predicting clinical adverse drug reactions (ADRs) of investigational drugs. We created the Secondary Pharmacology Database (SPD) by testing 1958 drugs using 200 assays to validate target-ADR associations. Compared to public and subscription resources, 95% of all and 36% of active (AC50 < 1 µM) results were unique to SPD, with bias towards higher activity in public resources. Annotating drugs with free maximal plasma concentrations, we found 684 physiologically relevant novel off-target activities. Furthermore, 64% of putative ADRs linked to target activity in key literature reviews were not statistically significant in SPD. Systematic analysis of all target-ADR pairs identified several novel associations confirmed by publications. Finally, candidate mechanisms for known ADRs are proposed based on SPD off-target activities. Taken together, we present a unique freely-available resource for benchmarking ADR predictions, explaining novel phenotypic activity and investigating clinical properties of marketed drugs.
These data constitute the supplementary material to the manuscript and are required as input for the Jupyter notebooks that contain the code for reproducing reported results.
Carbonic anhydrases (CAs) catalyze the physiological hydration of carbon dioxide and are among the most intensely studied pharmaceutical target enzymes. A hallmark of CA inhibition is the complexation of the catalytic zinc cation in the active site. Human (h) CA isoforms belonging to different families are implicated in a wide range of diseases and of very high interest for therapeutic intervention. Given the conserved catalytic mechanisms and high similarity of many hCA isoforms, a major challenge for CA-based therapy is achieving inhibitor selectivity for hCA isoforms that are associated with specific pathologies over other widely distributed isoforms such as hCA I or hCA II that are of critical relevance for the integrity of many physiological processes. To address this challenge, we have attempted to predict compounds that are selective for isoform hCA IX, which is a tumor-associated protein and implicated in metastasis, over hCA II on the basis of a carefully curated data set of selective and nonselective inhibitors. Machine learning achieved surprisingly high accuracy in predicting hCA IX-selective inhibitors. The results were further investigated, and compound features determining successful predictions were identified. These features were then studied on the basis of X-ray structures of hCA isoform-inhibitor complexes and found to include substructures that explain compound selectivity. Our findings lend credence to selectivity predictions and indicate that the machine learning models derived herein have considerable potential to aid in the identification of new hCA IX-selective compounds.
Aim: Providing compound data sets for promiscuity analysis with single-target (ST) and multi-target (MT) activity, taking confirmed inactivity against targets into account. Methodology: Compounds and target annotations are extracted from screening assays. For a given combination of targets, MT and ST compounds are identified, ensuring test data completeness. Exemplary results & data: A total of 1242 MT compounds active against five or more targets and 6629 corresponding ST compounds are characterized, organized and made freely available. Limitations & next steps: Screening campaigns typically cover a smaller target space than compounds from the medicinal chemistry literature and their activity annotations might be of lesser quality. Reported compound groups will be subjected to target set-based promiscuity analysis and predictions.
The compound optimization monitor (COMO) approach was originally developed as a diagnostic approach to aid in evaluating development stages of analog series and progress made during lead optimization. COMO uses virtual analog populations for the assessment of chemical saturation of analog series and has been further developed to bridge between optimization diagnostics and compound design. Herein, we discuss key methodological features of COMO in its scientific context and present a deep learning extension of COMO for generative molecular design, leading to the introduction of DeepCOMO. Applications on exemplary analog series are reported to illustrate the entire DeepCOMO repertoire, ranging from chemical saturation and structure–activity relationship progression diagnostics to the evaluation of different analog design strategies and prioritization of virtual candidates for optimization efforts, taking into account the development stage of individual analog series.
Small molecules with multitarget activity are capable of triggering polypharmacological effects and are of high interest in drug discovery. Compared to single-target compounds, promiscuity also affects drug distribution and pharmacodynamics and alters ADMET characteristics. Features distinguishing between compounds with single- and multitarget activity are currently only little understood. On the basis of systematic data analysis, we have assembled large sets of promiscuous compounds with activity against related or functionally distinct targets and the corresponding compounds with single-target activity. Machine learning predicted promiscuous compounds with surprisingly high accuracy. Molecular similarity analysis combined with control calculations under varying conditions revealed that accurate predictions were largely determined by structural nearest-neighbor relationships between compounds from different classes. We also found that large proportions of promiscuous compounds with activity against related or unrelated targets and corresponding single-target compounds formed analog series with distinct chemical space coverage, which further rationalized the predictions. Moreover, compounds with activity against proteins from functionally distinct classes were often active against unique targets that were not covered by other promiscuous compounds. The results of our analysis revealed that nearest-neighbor effects determined the prediction of promiscuous compounds and that preferential partitioning of compounds with single- and multitarget activity into structurally distinct analog series was responsible for such effects, hence providing a rationale for the presence of different structure-promiscuity relationships.
In medicinal chemistry, compound optimization largely depends on chemical knowledge, experience, and intuition, and progress in hit-to-lead and lead optimization projects is difficult to estimate. Accordingly, approaches are sought after that aid in assessing the odds of success with an optimization project and making decisions whether to continue or discontinue work on an analog series at a given stage. However, currently there are only very few approaches available that are capable of providing decision support. We introduce a computational methodology designed to combine the assessment of chemical saturation of analog series and structure-activity relationship (SAR) progression. The current endpoint of these development efforts, the compound optimization monitor (COMO), further extends lead optimization diagnostics to compound design and activity prediction. Hence, COMO plays dual role in supporting lead optimization campaigns.
Aim: Combining computational lead optimization diagnostics with analog design and computational approaches for assessing optimization efforts are discussed and the compound optimization monitor is introduced. Methods: Approaches for compound potency prediction are described and a new analog design algorithm is introduced. Calculation protocols are detailed. Results & discussion: The study rationale is explained. Compound optimization monitor diagnostics are combined with a thoroughly evaluated approach for compound design and candidate prioritization. The diagnostic scoring scheme is further extended. Future perspective: Opportunities for practical applications of the integrated computational methodology are described and further development perspectives are discussed.
Predicting compounds with single- and multi-target activity and exploring origins of compound specificity and promiscuity is of high interest for chemical biology and drug discovery. We present a large-scale analysis of compound promiscuity including two major components. First, high-confidence datasets of compounds with multi- and corresponding single-target activity were extracted from biological screening data. Positive and negative assay results were taken into account and data completeness was ensured. Second, these datasets were investigated using diagnostic machine learning to systematically distinguish between compounds with multi- and single-target activity. Models built on the basis of chemical structure consistently produced meaningful predictions. These findings provided evidence for the presence of structural features differentiating promiscuous and non-promiscuous compounds. Machine learning under varying conditions using modified datasets revealed a strong influence of nearest neighbor relationship on the predictions. Many multi-target compounds were found to be more similar to other multi-target compounds than single-target compounds and vice versa, which resulted in consistently accurate predictions. The results of our study confirm the presence of structural relationships that differentiate promiscuous and non-promiscuous compounds.
Future Science OAVol. 6, No. 8 CommentaryOpen AccessInhibitor bias in luciferase-based luminescence assaysDimitar Yonchev & Jürgen BajorathDimitar YonchevDepartment of Life Science Informatics, B-IT, LIMES Program Unit Chemical Biology & Medicinal Chemistry, Rheinische Friedrich-Wilhelms-Universität, Endenicher Allee 19c, D-53115 Bonn, Germany & Jürgen Bajorath *Author for correspondence: Tel.: +49 228 7369 100; Fax: +49 228 7369 101; E-mail Address: bajorath@bit.uni-bonn.dehttps://orcid.org/0000-0002-0557-5714Department of Life Science Informatics, B-IT, LIMES Program Unit Chemical Biology & Medicinal Chemistry, Rheinische Friedrich-Wilhelms-Universität, Endenicher Allee 19c, D-53115 Bonn, GermanyPublished Online:17 Jun 2020https://doi.org/10.2144/fsoa-2020-0081AboutSectionsPDF/EPUB ToolsAdd to favoritesDownload CitationsTrack Citations ShareShare onFacebookTwitterLinkedInRedditEmail Keywords: bioluminescence assayscomputational analysisfalse positivesfirefly luciferasehit ratesluciferase inhibitorsmechanisms of actionpublic assay dataThe firefly luciferase (FLuc)-based bioluminescent reaction [1] provides the basis for one of the most popular detection systems in high-throughput screening (HTS) [2], especially for cell-based (reporter gene) assays [2,3]. Bioluminescence results from the activity of FLuc, which catalyzes the ATP-dependent reaction of its natural substrate D-luciferin, a benzothiazole derivative, with oxygen that emits energy in the form of light [1]. Depending on the design of the detection system, an increase or decrease in the FLuc-dependent luminescence signal is measured. It is known that FLuc assay readouts are vulnerable to FLuc inhibition by small molecules [4,5]. FLuc inhibitors are frequently encountered and act by a variety of mechanisms including competitive, noncompetitive and anticompetitive (uncompetitive) inhibition [5]. Depending on the assay format, effects of direct FLuc inhibition can be complex and difficult to analyze. This especially applies when FLuc is used as a reporter in cell-based assays. FLuc is highly sensitive to proteolysis and only has a short half-life under cellular conditions [6]. A striking FLuc reporter assay interference effect that is counterintuitive at first sight is caused by the inhibition of FLuc, which actually results in an increase in the luminescence signal (rather than a decrease, as one might expect) [6–8]. In this case, inhibitor binding stabilizes the enzyme and protects it against degradation, which increases its half-life [6–8]. If inhibition retains baseline FLuc activity, the net effect of stabilization is an increase in the luminescence signal. Under assay conditions, increased emission of light through FLuc inhibition has been demonstrated to result from the formation of a multi-substrate adduct inhibitor (by an anticompetitive mechanism) [8]. For reporter gene assays relying on an increase in the luminescence signal relative to a control, such FLuc inhibition is highly likely to cause false positive assay readouts. In addition, various other inhibitory effects are possible, which potentially bias's FLuc-based assays.Herein, we have addressed the question how FLuc inhibition might affect different assays with FLuc-dependent readouts and if there may be general trends that can be detected. Therefore, a systematic computer-aided analysis of public screening data was carried out.FLuc inhibitors from public assay dataInitially, we collected experimentally confirmed FLuc inhibitors from the current PubChem BioAssay database [9], including assays that were imported from ChEMBL [10]. A set of 57 assays was obtained that aimed to identify FLuc inhibitors in different ways. These assays varied from large screens of more than 360,000 tested compounds (for example, PubChem assay ID: AID 588342), to very small assays containing only few compounds (for example, AID 2229) and often represented counter screens (such as AID 2515).Since assays for FLuc inhibitors carried out over time frequently yielded false negatives [5,11], FLuc inhibitors were selected for our analysis if they were classified as active at least once. To remove potential assay interference molecules from designated FLuc inhibitors, pan-assay interference compounds [12] and likely colloidal aggregators [13] were removed using computational filters. On the basis of these criteria, a total of 24,449 FLuc inhibitors were obtained. In addition, all compounds that were consistently inactive in FLuc assays were collected as FLuc noninhibitors.Assays with FLuc-dependent readoutsNext, we assembled (nonFLuc inhibitor) assays for other biological targets or phenotypes with FLuc-dependent readouts from the PubChem BioAssay database. For the identification of FLuc-based detection systems, information concerning the assay type, format and detection method were extracted from PubChem BioAssay records. In addition, key word searches were carried out in assay descriptions and protocols utilizing the string patterns 'lucife', 'lumin', 'glo' and 'ATPlite,' as additional indicators of relevant assays. Furthermore, assays with confirmed FLuc-dependent readouts were required to contain at least 100 tested compounds with 'active' or 'inactive' annotations, including at least one of the FLuc inhibitors we identified, as described above. On the basis of these selection criteria, 1014 FLuc technology-based assays were obtained, including HTS assays with more than 10,000 tested compounds. These assays contained one to 21,469 FLuc inhibitors, with a median value of 121 inhibitors per assay. They were predominantly cell-based and covered a wide spectrum of biological targets or phenotypes.Hit rate analysisIn FLuc-based assays, hit rates were separately determined per assay for all tested compounds and known FLuc inhibitors, respectively. Figure 1A compares these hit rates for different proportions of FLuc inhibitors among assay hits. The comparison reveals a clear trend of globally increasing assay hit rates for increasing hit rates of FLuc inhibitors. In other words, FLuc inhibitors were enriched among active compounds in assays with increasingly high hit rates. This finding indicated that FLuc inhibitors exhibited a general tendency to cause false positives in assays with FLuc-dependent detection systems, regardless of the assay format, the design of the detection system and the presence of increasing or decreasing luminescence signals as an indicator of activity. Similar effects were observed previously for a confined set of FLuc-based assays in which FLuc inhibition led to FLuc stabilization [7].Figure 1. Hit rate analysis.(A) Shown is a scatter plot of hit rates in FLuc-based assays (represented as dots). On the y-axis, hit rates per assay are reported for all tested compounds. On the x-axis, corresponding hit rates are reported for FLuc inhibitor subsets per assay. Dots are color-coded according to different proportions of FLuc inhibitors among assay hits (less than 25%, light blue; 25–50%, green; 50–75%, orange; 75% or more, red). The corresponding sections are separated by dashed gray lines. The pie chart on the lower right shows the percentage of assays falling into each of the four ranges. (B) Shown are boxplots comparing hit rates (y-axis) of known FLuc inhibitors (yellow) and noninhibitors (light green) in assays with FLuc-dependent (left) or -independent (right) detection systems.FLuc: Firefly luciferase.Based on our comparison of hit rates across all types of assays with FLuc-dependent readouts, the trend of FLuc inhibitors to cause false positive signals can be generalized. We further investigated this finding by comparing hit rates of FLuc inhibitors and FLuc noninhibitors across PubChem assays with or without FLuc-dependent readouts. The comparison is shown in Figure 1B. There was no detectable difference in the hit rate distribution of FLuc inhibitors and noninhibitors in assays with FLuc-independent detection systems. By contrast, in assays with FLuc-dependent readouts, there was a significant increase in the hit rate of FLuc inhibitors when compared with noninhibitors.Concluding discussionFLuc inhibition by small molecules is a potential caveat for HTS assays that employ FLuc-based luminescence detection. This particularly applies to cell-based assays using FLuc as a reporter. FLuc inhibition is mechanistically complex and has potential secondary effects. An important consequence is inhibitor-induced FLuc stabilization, which leads to an extension of its cellular half-life and to a net increase in the luminescence signal. Accordingly, for assays that rely on detecting an increase in the signal, the presence of some –but not all– FLuc inhibitors gives rise to false positive readouts, depending on their mode of action. Importantly, assays with FLuc-dependent readouts often have distinct formats and detection characteristics and may rely on increases or decreases of luminescence signals to identify active compounds. Even in the simplest scenario, considering competitive inhibition in the absence of secondary effects leading to a reduction of FLuc activity and resulting signals, luminescence assays may be affected in different ways, depending on their design. Therefore, we have been interested in investigating the question whether FLuc inhibition might result in general trends across assays with FLuc-dependent readouts having different formats and detection characteristics. To address this question, we have carried out the systematic analysis presented herein. Several observations are of note. FLuc inhibitors were demonstrated to be widely distributed across public domain assays, more so than we anticipated. This may be different for assays and curated source libraries used in the pharmaceutical industry. Regardless, if data from publicly available luminescence assays are used, for example, for generating screening statistics or deriving computational models for activity prediction, there should be awareness of potential sources of FLuc inhibitor bias and ensuing errors. However, the presence of FLuc inhibitors in large numbers of public assays enabled us to analyze their influence on hit rates on a large scale. Care was taken to omit known FLuc inhibitors from further consideration that were associated with other potential assay interference effects, thus focusing on FLuc inhibition as the major source of assay interference. Despite stringent selection, more than 24,000 qualifying FLuc inhibitors were obtained whose activity annotations across more than 1000 assays with FLuc-dependent readouts were examined. There was a general increase in hit rates of assays with FLuc-dependent readouts in the presence of increasing proportions of FLuc inhibitors, regardless of the assay types. This increase most likely resulted from altered luminescence signals and reflected a general tendency of subsets of FLuc inhibitors to become false positives in these assays. Hence, care must be taken when selecting active compounds from these assays for further studies. Clearly, one should be aware of likely FLuc inhibitor bias in FLuc-based luminescence assays given the frequency and magnitude of false positive contributions detected herein.AcknowledgmentsD Yonchev is supported by the Jürgen Manchot Stiftung.Financial & competing interests disclosureThe authors have no other relevant affiliations or financial involvement with any organization or entity with a financial interest in or financial conflict with the subject matter or materials discussed in the manuscript. This includes employment, consultancies, honoraria, stock ownership or options, expert testimony, grants or patents received or pending, or royalties.No writing assistance was utilized in the production of this manuscript.Open accessThis work is licensed under the Creative Commons Attribution 4.0 License. To view a copy of this license, visit http://creativecommons.org/licenses/by/4.0/Papers of special note have been highlighted as: • of interest; •• of considerable interestReferences1. Deluca M. Firefly luciferase. Adv. Enzymol. Relat. Areas Mol. Biol. 44(1), 37–68 (1976).CAS, Google Scholar2. Fan F, Wood KV. Bioluminescent assays for high-throughput screening. Assay Drug Dev. Technol. 5(1), 127–136 (2007).Crossref, CAS, Google Scholar3. An WF, Tolliday N. Cell-based assays for high-throughput screening. Mol. Biotechnol. 45(2), 180–186 (2010).Crossref, CAS, Google Scholar4. Braeuning A. Firefly luciferase inhibition: a widely neglected problem. Arch. Toxicol. 89(1), 141–142 (2015).Crossref, CAS, Google Scholar5. Thorne N, Shen M, Lea WA et al. Firefly luciferase in chemical biology: a compendium of inhibitors, mechanistic evaluation of chemotypes and suggested use as a reporter. Chem. Biol. 19(8), 1060–1072 (2012). •• In-depth analysis of a variety of firefly luciferase inhibitors and their modes of action.Crossref, CAS, Google Scholar6. Thompson JF, Hayes LS, Lloyd DB. Modulation of firefly luciferase stability and impact on studies of gene regulation. Gene 103(2), 171–177 (1991). • Refs. 6–8 detail a prominent mechanism of firefly luciferase cell-based assay interference.Crossref, CAS, Google Scholar7. Auld DS, Thorne N, Nguyen DT, Inglese J. A specific mechanism for nonspecific activation in reporter-gene assays. ACS Chem. Biol. 3(8), 463–470 (2008). • Prominent mechanism of firefly luciferasecell-based assay interference.Crossref, CAS, Google Scholar8. Auld DS, Lovell S, Thorne N et al. Molecular basis for the high-affinity binding and stabilization of firefly luciferase by PTC124. Proc. Nat. Acad. Sci. USA 107(11), 4878–4883 (2010). • Prominent mechanism of firefly luciferase cell-based assay interference.Crossref, CAS, Google Scholar9. Wang Y, Bryant SH, Cheng T et al. PubChem BioAssay: 2017 update. Nucleic Acids Res. 45(D1), D955–D963 (2017).Crossref, CAS, Google Scholar10. Gaulton A, Hersey A, Nowotka M et al. The ChEMBL database in 2017. Nucleic Acids Res. 45(D1), D945–D954 (2017).Crossref, CAS, Google Scholar11. Ghosh D, Koch U, Hadian K et al. Luciferase advisor: high-accuracy model to flag false positive hits in luciferase HTS assays. J. Chem. Inf. Model 58(5), 933–942 (2018).Crossref, CAS, Google Scholar12. Baell JB, Holloway GA. New substructure filters for removal of pan assay interference compounds (PAINS) from screening libraries and for their exclusion in bioassays. J. Med. Chem. 53(7), 2719–2740 (2010).Crossref, CAS, Google Scholar13. Irwin JJ, Duan D, Torosyan H et al. An aggregation advisor for ligand discovery. J. Med. Chem. 58(17), 7076–7087 (2015).Crossref, CAS, Google ScholarFiguresReferencesRelatedDetailsCited ByStructured data sets of compounds with multi-target and corresponding single-target activity from biological assaysChristian Feldmann, Dimitar Yonchev & Jürgen Bajorath11 March 2021 | Future Science OA, Vol. 7, No. 5Analysis of Biological Screening Compounds with Single- or Multi-Target Activity via Diagnostic Machine Learning27 November 2020 | Biomolecules, Vol. 10, No. 12 Vol. 6, No. 8 Follow us on social media for the latest updates Metrics Downloaded 1,158 times History Received 7 May 2020 Accepted 7 May 2020 Published online 17 June 2020 Published in print September 2020 Information© 2020 Jürgen BajorathKeywordsbioluminescence assayscomputational analysisfalse positivesfirefly luciferasehit ratesluciferase inhibitorsmechanisms of actionpublic assay dataAcknowledgmentsD Yonchev is supported by the Jürgen Manchot Stiftung.Financial & competing interests disclosureThe authors have no other relevant affiliations or financial involvement with any organization or entity with a financial interest in or financial conflict with the subject matter or materials discussed in the manuscript. This includes employment, consultancies, honoraria, stock ownership or options, expert testimony, grants or patents received or pending, or royalties.No writing assistance was utilized in the production of this manuscript.Open accessThis work is licensed under the Creative Commons Attribution 4.0 License. To view a copy of this license, visit http://creativecommons.org/licenses/by/4.0/PDF download
Aim: Development of a new, practically applicable computational method to monitor progress in lead optimization. Computational approaches that aid in compound optimization are discussed and the Compound Optimization Monitor (COMO) method is introduced and put into scientific context. Methodology & calculations: The methodological concept and the COMO scoring scheme are described in detail. Results & discussions: Calculation parameters are evaluated, and profiling results reported for an ensemble of analog series. Future perspective: The dual role of virtual analogs as diagnostic tools for progress evaluation and as potential candidates for lead optimization is discussed. In light of this dual role, interfacing COMO with machine learning for compound activity prediction and prioritization of candidates is highlighted as a future research objective.
Compounds with multitarget activity are of high interest for polypharmacological drug discovery. Such promiscuous compounds might be active against closely related target proteins from the same family or against distantly related or unrelated targets. Compounds with activity against distinct targets are not only of interest for polypharmacology but also to better understand how small molecules might form specific interactions in different binding site environments. We have aimed to identify compounds with activity against drug targets from different classes. To these ends, a systematic analysis of public biological screening data was carried out. Care was taken to exclude compounds from further consideration that were prone to experimental artifacts and false positive activity readouts. Extensively assayed compounds were identified and found to contain molecules that were consistently inactive in all assays, active against a single target, or promiscuous. The latter included more than 1000 compounds that were active against 10 or more targets from different classes. These multiclass ligands were further analyzed and exemplary compounds were found in X-ray structures of complexes with distinct targets. Our collection of multiclass ligands should be of interest for pharmaceutical applications and further exploration of binding characteristics at the molecular level. Therefore, these highly promiscuous compounds are made publicly available.
Public repositories of compounds and activity data are of prime importance for pharmaceutical research in academic and industrial settings. Major databases have evolved over the years. Their growth is accompanied by an increasing tendency toward data sharing. This is a positive development but not without potential problems. Using ChEMBL and PubChem as examples, we show that crosstalk between databases also leads to substantial data redundancy that might not be obvious. Redundancy is an important issue because it biases data analysis and knowledge extraction and leads to inflated views of available compounds, assays and activity data. Going forward it will be important to further refine data exchange and deposition criteria and make redundancy as transparent as possible.
Assessing the degree to which analogue series are chemically saturated is of major relevance in compound optimization. Decisions to continue or discontinue series are typically made on the basis of subjective judgment. Currently, only very few methods are available to aid in decision making. We further investigate and extend a computational concept to quantitatively assess the progression and chemical saturation of a series. To these ends, existing analogues and virtual candidates are compared in chemical space and compound neighborhoods are systematically analyzed. A large number of analogue series from different sources are studied, and alternative chemical space representations and virtual analogues of different designs are explored. Furthermore, evolving analogue series are distinguished computationally according to different saturation levels. Taken together, our findings provide a basis for practical applications of computational saturation analysis in compound optimization.
In medicinal chemistry, lead optimization is a critically important task and a highly empirical process, largely driven by chemical knowledge and intuition. Only very few approaches are available to guide and evaluate optimization efforts. It is often very difficult to understand when a compound series is exhausted and the generation of additional analogs unlikely to yield further progress toward potent and efficacious candidates. Rationalizing lead optimization remains an essentially unsolved problem. Herein, we introduce a new computational method to aid in evaluating whether sufficient numbers of analogs have been made and further progress is unlikely. The approach integrates the assessment of chemical saturation and structure-activity relationship progression of compound series. Easy-to-calculate scores characterize evolving analog series and identify candidates with high or low priority for further chemical exploration.
The secondary metabolism of bacteria, fungi and plants yields a vast number of bioactive substances. The constantly increasing amount of published genomic data provides the opportunity for an efficient identification of gene clusters by genome mining. Conversely, for many natural products with resolved structures, the encoding gene clusters have not been identified yet. Even though genome mining tools have become significantly more efficient in the identification of biosynthetic gene clusters, structural elucidation of the actual secondary metabolite is still challenging, especially due to as yet unpredictable post-modifications. Here, we introduce SeMPI, a web server providing a prediction and identification pipeline for natural products synthesized by polyketide synthases of type I modular. In order to limit the possible structures of PKS products and to include putative tailoring reactions, a structural comparison with annotated natural products was introduced. Furthermore, a benchmark was designed based on 40 gene clusters with annotated PKS products. The web server of the pipeline (SeMPI) is freely available at: http://www.pharmaceutical-bioinformatics.de/sempi.